Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific task…
arXiv API — AI safety & capability queryMeta AIretrieved 1 h ago
Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and the limitations of static embedding alignments. This study investigates the diachronic evolution of the Sinhala l…
arXiv API — AI safety & capability queryMeta AIretrieved 1 h ago
Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we compare linear probing,…
arXiv API — AI safety & capability queryMeta AIretrieved 1 h ago
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric cau…
arXiv API — AI safety & capability queryMeta AIretrieved 1 h ago