Tag: ai-safety
All the articles with the tag "ai-safety".
-
The Intelligent Worm: Adaptive Malware
A conceptual, defense-oriented threat analysis of the 'Intelligent Worm' — a hypothetical self-propagating program whose exploitation capability is regenerated by an AI reasoning loop rather than fixed at authoring time. I walk the history of static worms through an epidemiological lens (the reproductive number R0 and the epidemic threshold), show why patching has historically driven outbreaks below threshold, and argue that an adaptive exploit term changes the shape of that curve. I take the feasibility limits seriously — hallucination, verification cost, autonomous experimentation — survey what current LLM-agent security research (Morris II, the Fang et al. autonomous-exploitation results, Google's Big Sleep) actually demonstrates, and lay out the defensive re-architecture the threat model implies. No exploit code, no propagation implementation: this is a stress-test of our defensive assumptions.
-
Society of Models: A Citation-Grounded Survey of AI Agent Collaboration Research (2023–2026)
An end-to-end survey of AI agent collaboration as it stands in May 2026, grounded in 100+ citations from ICML/ICLR/NeurIPS 2023–2026 and arXiv. Architectures (AutoGen, MetaGPT, Magentic-One, OpenHands, ADAS, AFlow), communication protocols (MCP, A2A), critic/verifier patterns (Self-Refine, Reflexion, LLM-Modulo), planning and decomposition (ReAct, Tree-of-Thoughts, CodeAct, GPTSwarm), benchmarks (SWE-bench, GAIA, WebArena, OSWorld, τ-bench, MLE-bench, TheAgentCompany), the contrarian compute-fair-comparison literature (Stop Overvaluing MAD, Agentless, Mixture-of-Agents rethinks), and safety/control (AI Control, debate, collusion, AgentHarm). With negative results and the methodological crisis foregrounded.