Dera News
derafrom heavy users
17:18 JST
SATURDAY, 18 JULY 2026

half the internet is terrified of AI. we are on the other half, taking this into our daily life, trying to understand better.

we use AI in everything we do — so every monday we read the whole week of it and work out what actually happened. something real, from heavy users. it makes our day if what we produce makes someone find AI more interesting for their life.

read our version of what happened this week in AI. free.read this week →
01
Hugging Face Daily Papers5 HR AGO/ primary source

On Locality and Length Generalization in Visual Reasoning

Institution: Qualcomm | Authors: Pulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya arXiv Links arXiv | PDF AI summary Abstract A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular computer vision models in use today, which input images globally and in a single shot. A natural question the...

— the story beneath
02
Hugging Face Daily Papers7 HR AGO/ primary source

Rethinking the Evaluation of Harness Evolution for Agents

New research casts doubt on the standard 'automated harness evolution' method for evaluating LLM agents, citing risks of overfitting and limited real-

What happened
  • University of Washington researchers found 'automated harness evolution' for LLM agent evaluation doesn't consistently outperform simpler testing.
  • The method risks agents overfitting to specific benchmarks since the same tests are used for both refinement and final assessment.
Why it matters
  • This means current reported performance metrics for some LLM agents might be inflated or not reflect true capabilities beyond specific test sets.
  • It signals a need for the AI industry to develop more robust and unbiased evaluation standards to ensure agent reliability and generalizability.
03
Hugging Face Daily Papers15 HR AGO/ primary source

Token Time Continuous Diffusion for Language Modeling

Institution: University of Texas at Austin | Authors: Parikshit Bansal, Sujay Sanghavi arXiv Links arXiv | PDF AI summary Abstract In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates in continuous space, deterministically mapping Gaussian noise to a final token canvas with no further sampling, and crucially (b) incorporates a new notion of per-token times, with some tokens proceeding from noise to token at a faster rate than others...

04
MarkTechPost21 MIN AGO

Google Cloud’s Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite

Google Cloud's generative-ai repository ships the Always-On Memory Agent, a reference implementation that treats memory as a running process. Built on Google ADK and Gemini 3.1 Flash-Lite, it uses no vector database and no embeddings. Instead, an orchestrator routes to Ingest, Consolidate, and Query sub-agents that read, connect, and write structured memory into SQLite 24/7. The post Google Cloud’s Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1...

05
GitHub Blog16 HR AGO/ primary source

The cost of saying yes has changed

The cost of writing code dropped; the cost of owning it didn't. A framework for deciding which changes are actually cheap in the AI era. The post The cost of saying yes has changed appeared first on The GitHub Blog.

06
VentureBeat14 HR AGO

Brex built its AI agent policy by watching what agents actually do, not by writing rules first

OpenClaw has become one of the most widely adopted agentic frameworks, but it has yet to prove itself at enterprise scale. Agents need real credentials — API keys, OAuth tokens, service accounts — to work effectively, and Brex found that traditional guardrails couldn't contain what those agents were doing with them. Brex set out to overcome these limitations by building an internal platform it calls CrabTrap. The open-source HTTP/HTTPS proxy intercepts all network traffic, examines policy rules,...

— the rundown
07

Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do

Capital One on Thursday released VulnHunter, an open-source, agentic AI security tool that scans source code for exploitable vulnerabilities, maps out how an attacker would reach them, and proposes targeted fixes — all before a single line ships to production. The tool, built internally and now available on GitHub under an Apache 2.0 license, is one of the most ambitious attempts by a major financial institution to turn offensive AI capabilities into a public defensive resource. The move marks a...

VentureBeat11 HR AGO
10

BadWAM: When World-Action Models Dream Right but Act Wrong

Authors: Qi Li, Xingyi Yang, Xinchao Wang arXiv Links arXiv | PDF AI summary Abstract World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assu...

Hugging Face Daily Papers29 HR AGO/ primary source
11

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Authors: Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu arXiv Links arXiv | PDF AI summary Abstract Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predic...

Hugging Face Daily Papers15 HR AGO/ primary source
12

Hierarchical Denoising For Multi-Step Visual Reasoning

Authors: Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan arXiv Links arXiv | PDF AI summary Abstract Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradigms struggle to achieve logical consistency and low-latency s...

Hugging Face Daily Papers15 HR AGO/ primary source
14

Sakana AI’s Error Diffusion Trains Dale-Compliant Dual-Stream Networks, Reaching 96.7% MNIST and 61.7% CIFAR-10 Without Backpropagation

Backpropagation relies on weight transport, which biological circuits likely cannot implement. Sakana AI's Error Diffusion sidesteps that constraint, training dual-stream excitatory/inhibitory networks that obey Dale's principle. This piece breaks down how modulo error routing scales the rule from MNIST to CIFAR-10 and reinforcement learning, and what its task-dependent ablations reveal. The post Sakana AI’s Error Diffusion Trains Dale-Compliant Dual-Stream Networks, Reaching 96.7% MNIST and 61....

MarkTechPost2 HR AGO
16

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

Institution: Mind Lab | Authors: Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao arXiv Links arXiv | PDF AI summary Abstract A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment. The gap is especially important for AI agents, whose observations, tool outputs, documents, and prior decisions accumu...

Hugging Face Daily Papers21 HR AGO/ primary source
19

Apple’s plot to crush OpenAI

Apple has reportedly filed a lawsuit against OpenAI, raising questions about its AI strategy and competitive landscape.

The Verge15 HR AGO
786 stories scored · 14-day window · 9/20 fully briefed/ranked entirely by dera's own scoring · the score stays hidden

Weekly AI brief for practitioners

A practitioner-first AI newsletter by dera.ai — what changed, why it matters, what to try next

By dera.ai • Every Monday, free

We respect your privacy. Unsubscribe at any time.