Resilient and Safe AI Systems
Graceful degradation under failures, and validation of AI-generated decisions before they reach the control loop
Classical and Byzantine fault tolerance mask faults through replication and voting, but the resource-constrained edge cannot afford full replication. Voting also fails for AI, because identical model replicas tend to fail on the same inputs in the same way. My research keeps AI systems dependable in two ways, by staying available when infrastructure fails and by staying safe when AI itself generates the decisions that drive the system.
Graceful Degradation under Failure
The edge cannot afford full replication, the classic recipe for resilience that keeps a spare copy of every model. FailLite (Wu et al., 2025) makes model serving failure-resilient without that cost through heterogeneous replication. Instead of identical copies, it backs each model with cheaper, smaller variants, so a failed serving node degrades to a lighter model rather than dropping requests. This builds on an earlier result, that falling back to a smaller model through graceful service degradation (Hanafy et al., 2023) beats losing the request outright.
The same principle extends beyond single models. It reaches the practical realities of deploying resilient ML at the edge (Practical Considerations (Gudipaty et al., 2025)) and full multi-stage inference pipelines rather than a single model (Resilient Distributed ML Inference Pipelines (Wu et al., 2024)).
Ensembles for Resilience
Heterogeneous backups help most when the models fail differently from one another. MEL (Gudipaty et al., 2025) trains a set of lightweight models that refine each other’s predictions when multiple servers are available, and each model still performs well on its own when failures cut the ensemble down to one. A diversity-encouraging training objective makes the models complementary rather than redundant, so the ensemble stays accurate as servers drop out.
Validating AI-Generated Decisions
Resilience keeps a system running. Safety keeps it from acting on bad decisions. As devices increasingly use large language models to configure and coordinate with unfamiliar peers on the fly, LLM output cannot be trusted blindly in a control loop.
CollabIoT (Shastri et al., 2025) makes LLM-driven configuration safe for transient IoT collaboration. It uses an LLM to translate a user’s high-level intent into fine-grained, capability-based access-control policies, validates those policies before they take effect, and confines LLM reasoning to a control plane while lightweight proxies enforce only validated decisions in a low-latency data plane.
This approach is part of our broader vision, the Internet of Collaborating Things (Hanafy et al., 2026), in which agentic edge AI negotiates and executes cross-domain collaborations while keeping unsafe or incorrect AI-generated decisions out of the control loop.
References
2026
- Under ReviewThe Internet of Collaborating Things: Agentic Edge AI for Autonomous Cross-Domain CollaborationsDec 2026