Signals are promoted only when they change a judgment, strengthen a convergence, or create a time-sensitive advisory move.
Most recent successful retrieval (partial coverage): 01 Oct, 14:41 Beirut Packet generated 01 Oct, 14:41 Beirut Page built 01 October 2026, 14:41 Beirut
1do now
9open decisions
12signals in view
12precursors
72sources tracked
4stale · 4 errors
Today
Recent source-dated developments; a signal alone does not change advice
PUBLISHED01 Oct, 07:00
MOBA-VL: Event-Localized Multi-Turn Reinforcement Learning for Real-Time MOBA Commentary
preprint; not peer reviewed
arXiv:2609.38428v1 Announce Type: new Abstract: Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and…
Why you may careCould matter to agent workflows if the reported task success, human handoff, cost, and operating limits hold up.
Next checkCheck the measured outcome, comparison, limits, and whether it applies to a decision you own. Until then, treat it as a lead.
PUBLISHED01 Oct, 07:00
VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation
preprint; not peer reviewed
arXiv:2609.38512v1 Announce Type: new Abstract: Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four…
Why you may careCould matter to agent workflows if the reported task success, human handoff, cost, and operating limits hold up.
Next checkCheck the measured outcome, comparison, limits, and whether it applies to a decision you own. Until then, treat it as a lead.
PUBLISHED01 Oct, 07:00
RoboAssist: Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance
preprint; not peer reviewed
arXiv:2609.39384v1 Announce Type: new Abstract: Long-horizon surgical assistance requires humanoid robots to coordinate with evolving human activities while maintaining safety across planning and execution. We present RoboAssist, an agent-based framework for interactive human-humanoid planning that integrates…
Why you may careCould matter to agent workflows if the reported task success, human handoff, cost, and operating limits hold up.
Next checkCheck the measured outcome, comparison, limits, and whether it applies to a decision you own. Until then, treat it as a lead.
WATCHDECISION PROPOSED
Watch VAmoS as a preprint benchmark lead; do not change production routing or infer real-world voice-agent reliability from its reported simulations.
The authors report completion from 17.3% to 44.7% across 14 voice stacks and a large drop with background TV in a 100-call simulated utility-billing benchmark. These claims highlight a testable reliability constraint but remain unreviewed and unreplicated.
Say this to MJ agent workflows and voice-agent deployment adviceUse the benchmark as a prompt to test multi-request completion, background speech, action correctness and caller verification on a fixed representative task set. The preprint alone does not establish production failure rates.
Next moveMonitor for released benchmark materials and an independent reproduction; if available, run one bounded task comparison that records success, corrections and elapsed time.
What would change this judgment?
Independent reproduction shows materially higher completion under comparable multi-request and background-speech conditions, or the simulation fails to transfer to relevant real-call tasks.
Labs, models & tooling33/34 live
97%
Research13/13 live
100%
Expert intelligence6/6 live
100%
Policy, capital & adoption14/17 live
82%
Talent2/2 live
100%
Grok / X precursor scoutLatest attempt 24 Sep, 10:53 Beirut
Product Hunt AI launches: Client error '403 Forbidden' for url 'https://www.producthunt.com/categories/ai-agents'
For more information check: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/403
Congress — frontier technology bills and actions: credential required: CONGRESS_API_KEY
Regulations.gov — AI documents and comment deadlines: credential required: REGULATIONS_API_KEY
SAM.gov — AI procurement notices: credential required: SAM_API_KEY
Source cadence and retrieval status
Source
Cadence
Last success
Next due
Current state
OpenAI official news
30 min
01 Oct, 14:41
01 Oct, 15:11
ok
Anthropic official news
30 min
01 Oct, 14:41
01 Oct, 15:11
ok
Google DeepMind official blog
30 min
01 Oct, 14:41
01 Oct, 15:11
ok
Google Research blog
30 min
01 Oct, 14:41
01 Oct, 15:11
ok
Hugging Face blog
120 min
01 Oct, 13:41
01 Oct, 15:41
unchanged
NVIDIA technical blog
120 min
01 Oct, 14:11
01 Oct, 16:11
unchanged
arXiv AI preprints
360 min
01 Oct, 13:41
01 Oct, 19:41
unchanged
arXiv robotics preprints
360 min
01 Oct, 13:41
01 Oct, 19:41
unchanged
arXiv machine learning preprints
360 min
01 Oct, 13:41
01 Oct, 19:41
unchanged
vLLM releases
60 min
01 Oct, 14:11
01 Oct, 15:11
unchanged
llama.cpp releases
60 min
01 Oct, 14:11
01 Oct, 15:11
ok
Transformers releases
60 min
01 Oct, 14:11
01 Oct, 15:11
ok
SGLang releases
60 min
01 Oct, 14:11
01 Oct, 15:11
unchanged
Qwen model repositories
60 min
01 Oct, 14:11
01 Oct, 15:11
ok
deepseek-ai model repositories
60 min
01 Oct, 14:11
01 Oct, 15:11
ok
meta-llama model repositories
60 min
01 Oct, 14:11
01 Oct, 15:11
ok
Epoch AI — Gradient Updates
120 min
01 Oct, 14:11
01 Oct, 16:11
ok
METR research
120 min
01 Oct, 14:11
01 Oct, 16:11
unchanged
The Innermost Loop
30 min
01 Oct, 14:41
01 Oct, 15:11
ok
Moonshots with Peter Diamandis
30 min
01 Oct, 14:41
01 Oct, 15:11
unchanged
Jack Clark — Import AI
180 min
01 Oct, 13:11
01 Oct, 16:11
ok
arXiv computational linguistics preprints
360 min
01 Oct, 14:11
01 Oct, 20:11
unchanged
arXiv computer vision preprints
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
arXiv quantitative biology preprints
360 min
01 Oct, 14:11
01 Oct, 20:11
unchanged
Hugging Face papers
360 min
01 Oct, 14:11
01 Oct, 20:11
ok
Hugging Face recent model repositories
60 min
01 Oct, 14:11
01 Oct, 15:11
ok
Hugging Face recent dataset repositories
120 min
01 Oct, 14:11
01 Oct, 16:11
ok
GitHub trending repositories
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
GitHub trending Python repositories
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
GitHub trending Jupyter repositories
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
OpenRouter newest model listings
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
SWE-bench leaderboard
360 min
01 Oct, 08:41
01 Oct, 14:41
unchanged
Product Hunt AI launches
360 min
30 Sep, 07:01
01 Oct, 16:41
error · stale
SEC latest EDGAR filings
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
ARPA-E funding opportunities
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
ARPA-H funding opportunities
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
DARPA opportunities
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
FDA press announcements
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
NIST AI publications
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
LessWrong latest posts
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Hacker News newest
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Alignment Forum latest posts
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Anthropic careers
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Google DeepMind careers
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
GitHub topic artificial intelligence
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
GitHub topic agents
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
GitHub topic robotics
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Hugging Face trending Spaces
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
OpenRouter monthly rankings
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Artificial Analysis models
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
FDA AI/ML-enabled medical devices
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
NIST AI Resource Center
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
EU AI Office
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Defense Innovation Unit open solicitations
360 min
01 Oct, 08:41
01 Oct, 14:41
unchanged
DOE EERE funding opportunities
360 min
01 Oct, 08:41
01 Oct, 14:41
unchanged
AI Engineer events
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Foresight Institute events
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Copenhagen Institute for Futures Studies
360 min
01 Oct, 08:41
01 Oct, 14:41
unchanged
SynBioBeta
360 min
01 Oct, 08:41
01 Oct, 14:41
unchanged
NeurIPS program site
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
ICML program site
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
ICLR program site
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
bioRxiv AI-relevant biology preprints
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
medRxiv computational and clinical preprints
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
ClinicalTrials.gov AI studies
360 min
01 Oct, 14:11
01 Oct, 20:11
unchanged
GitHub trending developers
360 min
01 Oct, 08:41
01 Oct, 14:41
ok
Federal Register — artificial intelligence
30 min
01 Oct, 14:11
01 Oct, 14:41
ok
Federal Register — compute, data centers and energy
30 min
01 Oct, 14:11
01 Oct, 14:41
ok
Federal Register — biotechnology and clinical AI
60 min
01 Oct, 14:11
01 Oct, 15:11
ok
Congress — frontier technology bills and actions
60 min
Unknown
01 Oct, 23:41
error · stale
Regulations.gov — AI documents and comment deadlines
60 min
Unknown
01 Oct, 23:41
error · stale
SAM.gov — AI procurement notices
360 min
Unknown
01 Oct, 23:41
error · stale
Early observations
Unverified candidates for review. Ranking is a heuristic, not confidence or probability.
{
"abstract": "NHTSA, on behalf of the U.S. Department of Transportation (DOT), is substantially recalibrating the Corporate Average Fuel Economy (CAFE) program to bring the program into compliance with the law and to remove previous regulatory distortions which have induced manufacturers to make design decisions that have…
capabilityconstraintspolicy
Federal Register document metadata and abstract; underlying action requires review
{
"abstract": "NHTSA, on behalf of the U.S. Department of Transportation (DOT), is substantially recalibrating the Corporate Average Fuel Economy (CAFE) program to bring the program into compliance with the law and to remove previous regulatory distortions which have induced manufacturers to make design decisions that have…
capabilityconstraintspolicy
Federal Register document metadata and abstract; underlying action requires review
arXiv:2609.38428v1 Announce Type: new Abstract: Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives.…
arXiv:2609.38512v1 Announce Type: new Abstract: Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four requests.…
arXiv:2609.39384v1 Announce Type: new Abstract: Long-horizon surgical assistance requires humanoid robots to coordinate with evolving human activities while maintaining safety across planning and execution. We present RoboAssist, an agent-based framework for interactive human-humanoid planning that integrates workflow…
Accurate classification of lung adenocarcinoma growth patterns from hematoxylin and eosin (H&E)-stained whole-slide images is essential for treatment planning and prognosis. We develop and evaluate an exponentially weighted ensemble deep learning approach that combines five architectures (EfficientNet-B3, DeiT3, Swin…
capabilitydeploymentconstraints
medRxiv preprint; not peer reviewed and not clinical advice
arXiv:2609.31631v2 Announce Type: new Abstract: Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts.…
arXiv:2609.31635v1 Announce Type: new Abstract: Clinical trajectory models are usually evaluated by next-event accuracy on observed histories. Simulation is different: models must condition on their own generated events, allowing errors to compound. Although this problem is well known in sequence modelling, it has not been…
arXiv:2609.31669v1 Announce Type: new Abstract: We present NanoForecast v0.5, a 6.5M-parameter forecaster that competes with models 31x its size (TimesFM, 200M parameters) after training pipeline fixes and no architecture change. Retraining the v0.3 architecture with corrected loss-scope handling, tensor shape alignment, and…
arXiv:2609.32083v1 Announce Type: new Abstract: Deploying AI systems requires deciding when to delegate tasks and when humans should intervene to monitor and mitigate risk induced by AI operations. These decisions are challenging when failures cluster: a hallucination or harmful output can trigger further errors, creating…
arXiv:2609.32267v1 Announce Type: new Abstract: Releasing a model update requires certifying that its current-population risk stays below a threshold. Trusted labels are expensive, while a cheap evaluator, such as an LLM judge, scores every example. Reusing evaluator errors from earlier audits is tempting, but when may such…
arXiv:2609.32302v1 Announce Type: new Abstract: Prompt engineering has become an indispensable tool for using large language models (LLMs), turning LLMs into task-specific experts without changing their weights. Despite notable theoretical advances in prompt engineering, the theory for the more practical hard or discrete…
capabilitydeploymentconstraints
preprint; not peer reviewed
Reviewed developments
Proposed decisions, not executed actions. Breakthrough verification is not yet established.
DECISION PROPOSED · 17 Sep, 08:20
Upgrade llama.cpp before the next Apple Silicon MoE evaluation.
Decision rationale; primary evidence not yet reviewed.
Recorded rationaleThe release fixes a Metal path that can turn large activations into NaNs. Any earlier MoE result on that path may be unreliable.
Reviewed evidenceDecision rationale; primary evidence not yet reviewed.
DECISION PROPOSED · 17 Sep, 08:20
Do not adopt Mimir 1B from parameter count alone.
Decision rationale; primary evidence not yet reviewed.
Recorded rationalellama.cpp reports roughly four times the decode work of an equal-width dense model and about 3 GB of F16 KV cache at 4K context.
Reviewed evidenceDecision rationale; primary evidence not yet reviewed.
DECISION PROPOSED · 17 Sep, 08:20
Treat Groq 3 LPX power-efficiency claims as a future infrastructure signal, not available GCC capacity.
Decision rationale; primary evidence not yet reviewed.
Recorded rationaleThe vendor framing strengthens the case that power per useful token is the bottleneck, but it provides no dated GCC commissioning or customer-access evidence.
Reviewed evidenceDecision rationale; primary evidence not yet reviewed.
Decision queue
Every item has a next move and a condition that can overturn it.
DO NOWDECISION PROPOSED
Upgrade llama.cpp before the next Apple Silicon MoE evaluation.
The release fixes a Metal path that can turn large activations into NaNs. Any earlier MoE result on that path may be unreliable.
Say this to AI implementation teams running local modelsPause comparisons made on the affected Apple Silicon path. Re-run one representative task on the fixed build before treating earlier quality or latency results as decision-grade.
Next movePin a build at or after b10994, rerun one representative MoE prompt, and compare output plus latency with the previous build.
What would change this judgment?
The affected mul_mm_id path was not used by our model or backend configuration.
WATCHDECISION PROPOSED
Do not adopt Mimir 1B from parameter count alone.
llama.cpp reports roughly four times the decode work of an equal-width dense model and about 3 GB of F16 KV cache at 4K context.
Say this to CIOs and private-AI operatorsA small parameter count is not a low operating cost. Require cost per successful task, memory use and supervision time before choosing this model for private workflows.
Next moveWait for prefix-LM support and an independent cost-per-success result before spending evaluation time.
What would change this judgment?
A reproducible task test shows enough accuracy gain to offset the reported memory and decode costs.
WATCHDECISION PROPOSED
Treat Groq 3 LPX power-efficiency claims as a future infrastructure signal, not available GCC capacity.
The vendor framing strengthens the case that power per useful token is the bottleneck, but it provides no dated GCC commissioning or customer-access evidence.
Say this to GCC sovereign, infrastructure and data-center leadersPlan around power per useful AI outcome and verified access dates. Do not count announced silicon as sovereign capacity until an operating deployment exposes customer access and independent performance data.
Next moveTrack independent watts-per-token and customer-access evidence tied to an operating Vera Rubin deployment.
What would change this judgment?
A dated operational deployment exposes customer access and independently measured performance per watt.
WATCHDECISION PROPOSED
Watch Anthropic ART as a research-automation case; do not treat as a validated application.
The described pipeline links parallel genomic search to candidate triage and human lab testing, but ART function remains unknown and independent validation is absent.
Say this to USEK research and AI workflow advisorsTreat this as a workflow signal: agents can search and rank candidates, while human scientists retain experimental validation. Do not infer a usable biological tool or general autonomous discovery from this case.
Next moveReview the technical report and seek independent replication; measure cost, expert review time, candidate yield and repeatability.
What would change this judgment?
The technical report or independent replication fails to confirm the reported RNA expression or novelty, or comparable searches fail to reproduce the candidate workflow.
WATCHDECISION PROPOSED
Watch Ringg's production pattern; do not treat the customer-story performance figures as independently validated.
OpenAI describes task-based model routing, specialist agents, offline evaluation, staged rollout, live endpoint monitoring and human escalation in a deployed multi-channel service workflow.
Say this to MJ's agent workflows, USEK applied research, and GCC service deploymentsFor an agent deployment, pair task routing and specialist steps with offline evaluation, gradual release, endpoint health checks and clear human handoff. Treat the reported resolution, CSAT and savings figures as vendor/customer claims until independently reproduced.
Next moveAsk for independently measured task mix, resolution denominator, escalation rate, language-specific quality, cost per completed task and post-deployment monitoring evidence.
What would change this judgment?
Independent customer or auditor data shows that completion rates, quality, or cost improvements do not generalize beyond selected workloads or fail to include human handling costs.
WATCHDECISION PROPOSED
Track AnewDDE as a research-automation signal; do not recommend adoption from the preprint claim alone.
A bioRxiv preprint describes an agentic closed-loop drug-discovery workflow connecting structure, affinity, design and experiment selection. It reports 10.7% success for single-digit-nanomolar binders in one nanobody campaign; this remains author-reported evidence.
Say this to USEK research automation and MJ agent workflowsPotentially relevant to a bounded literature or lab-workflow review if methods and data substantiate the claim; no deployment decision yet.
Next moveInspect full methods and results, especially campaign denominator, assay protocol, controls, and availability of data/code.
What would change this judgment?
Full methods fail to support the reported binder yield, the result is not reproducible, or performance depends on a selected campaign that does not generalize.
WATCHDECISION PROPOSED
Use claude-opus-5-5 for approved premium Claude work after confirming the exact provider model ID; evaluate task-level quality, time, and spend before broadening its role.
Anthropic announced Opus 5.5 on Sep 22 and claims 40% lower typical token-billed running cost than Opus 5. The official model page gives the API identifier claude-opus-5-5. Its quality and cost claims remain vendor-reported.
Say this to MJ agent workflows and Claude routingPin claude-opus-5-5 for eligible premium tasks only after checking the active provider supports that exact ID; record outcome and measured usage against a fixed task baseline.
Next moveVerify exact provider model ID and run one bounded representative task comparison against the current Opus baseline, recording quality, elapsed time, token use and billed cost.
What would change this judgment?
Provider does not expose the exact model ID, or a task-matched comparison shows worse quality or no worthwhile end-to-end cost/time benefit.
WATCHDECISION PROPOSED
Treat OpenRouter's DeepSeek Latest price/catalog revision as a route-specific cost signal; verify an exact pinned model/version before updating any cost assumption or routing decision.
OpenRouter's current pages map Pro Latest to V4 Pro 0813 at $3.78/M output tokens and Flash Latest to V4.1 Flash at $0.60/M. The prior catalog snapshot differs materially, but latest aliases and slight page/API discrepancies prevent a like-for-like tariff conclusion.
Say this to MJ agent workflows, USEK research, and GCC deployment cost planningBefore relying on a DeepSeek Latest route for cost estimates, pin the exact model version and check the current provider/router tariff against the earlier catalog snapshot. No routing change is proposed from this listing alone.
Next moveCompare the pinned V4 Pro 0813 and V4.1 Flash tariffs with any direct-provider price sheet and calculate cost per successful representative task before changing a budget or route.
What would change this judgment?
The apparent increase disappears when comparing the same pinned model version, or direct provider pricing and a fixed-task cost-per-success measurement show no meaningful change for MJ's workload.
WATCHDECISION PROPOSED
Watch VAmoS as a preprint benchmark lead; do not change production routing or infer real-world voice-agent reliability from its reported simulations.
The authors report completion from 17.3% to 44.7% across 14 voice stacks and a large drop with background TV in a 100-call simulated utility-billing benchmark. These claims highlight a testable reliability constraint but remain unreviewed and unreplicated.
Say this to MJ agent workflows and voice-agent deployment adviceUse the benchmark as a prompt to test multi-request completion, background speech, action correctness and caller verification on a fixed representative task set. The preprint alone does not establish production failure rates.
Next moveMonitor for released benchmark materials and an independent reproduction; if available, run one bounded task comparison that records success, corrections and elapsed time.
What would change this judgment?
Independent reproduction shows materially higher completion under comparable multi-request and background-speech conditions, or the simulation fails to transfer to relevant real-call tasks.
Lead / lag scoreboard
63 matched items. Baseline discoveries are shown but never claimed as wins.
VERIFIED LEAD105h ahead
Self-Play Pretraining with Zero Data
Compared with Welcome to September 29, 2026. Prospective timing is measurable.
VERIFIED LEAD80h ahead
Introducing MentalHealthBench
Compared with Welcome to September 27, 2026. Prospective timing is measurable.
VERIFIED LEAD80h ahead
Contrastive Language Models
Compared with Welcome to September 27, 2026. Prospective timing is measurable.
VERIFIED LEAD59h ahead
Promising discoveries about the potential for life on one of Saturn’s icy moons
Compared with Welcome to September 29, 2026. Prospective timing is measurable.
VERIFIED LEAD43h ahead
Introducing Astra for Law
Compared with Welcome to September 20, 2026. Prospective timing is measurable.
VERIFIED LEAD43h ahead
Sep 18, 2026 Announcements Partnering with Accenture on embedded evaluation
Compared with Welcome to September 20, 2026. Prospective timing is measurable.
arXiv:2609.38428v1 Announce Type: new Abstract: Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second, both fluently and accurately. Existing streaming VLMs sound natural but often miss key events such as kills and objectives.…
arXiv:2609.38512v1 Announce Type: new Abstract: Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four requests.…
arXiv:2609.39384v1 Announce Type: new Abstract: Long-horizon surgical assistance requires humanoid robots to coordinate with evolving human activities while maintaining safety across planning and execution. We present RoboAssist, an agent-based framework for interactive human-humanoid planning that integrates workflow…
arXiv:2609.35882v1 Announce Type: cross Abstract: Genome foundation models are most useful where sequences are generated, yet the largest models need datacenter accelerators and a place to send private DNA. We ask whether a 15-billion-parameter genome mixture-of-experts (MoE) model can instead run on volunteers' web…
arXiv:2609.37456v1 Announce Type: new Abstract: Robotic ultrasound commonly targets standardized views, predefined scanning protocols, or expert-scan reproduction. We target a different role: while a clinician performs the primary procedure, a robotic assistant maintains a clinician-selected view so that changing tissue…
"Intelligence too cheap to meter" is a dream peddled by AI companies and accelerationists. In a narrow sense it is already true: if intelligence means solving coding and mathematics problems, AI is cheap today. But anyone who has spent much time with current models knows they cannot yet replicate human remote work. They are…
TL;DR We seek to identify open source reasoning models which are both small enough for white-box interpretability and display evaluation-gaming behavior. We find that how often models verbalize their awareness varies from very rarely to a third of the time, and appears largely unrelated to model size. Within the same…
arXiv:2609.38485v1 Announce Type: new Abstract: Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known to interfere. Existing diagnoses and remedies operate at the resolution of layers or experts, measuring conflict per layer and…
arXiv:2609.38758v1 Announce Type: new Abstract: Referring Video Object Segmentation (RVOS) aims to produce a pixel-accurate mask sequence for an object specified by natural language. Sa2VA combines a multimodal large language model with SAM2 for grounded segmentation; however, its inference typically grounds the query from a…
arXiv:2609.38811v1 Announce Type: new Abstract: Metal additive manufacturing parts are inspected by X-ray computed tomography, where labelled data is scarce, the pores and inclusions that matter span a few pixels, and inspection must happen at the machine. We present DCM-SAM, a defect-conditioned adaptive mixture of LoRA…
arXiv:2609.35794v1 Announce Type: new Abstract: Just as Socrates recognized the limits of his own knowledge, Retrieval-Augmented Language Models (RALMs) should learn to abstain when the retrieved evidence cannot support a reliable response. Existing approaches largely rely on monolithic LLMs to handle heterogeneous retrieval…
arXiv:2609.35815v1 Announce Type: new Abstract: Researchers across academia increasingly base significance claims on LLM judge scores and small-sample AI evaluations. Yet without well-calibrated confidence intervals (CIs), hypothesis tests, and judge-bias corrections, such claims are unreliable. We address these issues in…
capabilitydeploymentconstraints
preprint; not peer reviewed
Hypothesis register
Amber means the prerequisite still lacks reviewed evidence.
agent-reliability11 days overdue
When can agents complete our multi-step work with fewer interventions?
Reliable extended task completion5 reviewed
Lower human correction time1 reviewed
Independent task reproduction0 reviewed
Next discriminating observationIndependent fixed-task results reporting failures, retries and human minutes
open-model-economics11 days overdue
When do deployable open models become viable for our private workflows?
Usable license and released weights1 reviewed
Fits available memory1 reviewed
Acceptable quality at fully loaded cost1 reviewed
Next discriminating observationA reproducible cost-per-success comparison under our hardware constraints
gcc-compute-bottlenecks11 days overdue
Does announced sovereign compute translate into usable capacity?
Delivered equipment0 reviewed
Commissioned power1 reviewed
Accessible operational service0 reviewed
Next discriminating observationDated commissioning and customer-access evidence, not another capacity pledge
research-automation11 days overdue
Which scientific workflows now produce independently validated results?
Reliable execution0 reviewed
Affordable validated result0 reviewed
Experimental access2 reviewed
Independent scientific validation0 reviewed
Repeat adoption0 reviewed
Next discriminating observationIndependent replication including total cost and expert verification time
eval-credibility11 days overdue
Which frontier capability claims survive independent evaluation and provenance scrutiny?
Evaluator independence disclosed0 reviewed
Task and harness equivalence established0 reviewed
Independent reproduction0 reviewed
Contamination and privileged-access risks addressed0 reviewed
Next discriminating observationA primary evaluation artifact and an independent reproduction using the same task definition
This page is a decision surface, not a feed reader. “Decision proposed” records an advisory recommendation; it does not prove execution. Repeated coverage does not count as independent evidence. Unknown measurements remain unknown. Email inventory is incomplete; this page does not represent a complete account inventory.