Resilient and Safe AI Systems

Graceful degradation under failures, and validation of AI-generated decisions before they reach the control loop

Classical and Byzantine fault tolerance mask faults through replication and voting, but the resource-constrained edge cannot afford full replication. Voting also fails for AI, because identical model replicas tend to fail on the same inputs in the same way. My research keeps AI systems dependable in two ways, by staying available when infrastructure fails and by staying safe when AI itself generates the decisions that drive the system.


Graceful Degradation under Failure

The edge cannot afford full replication, the classic recipe for resilience that keeps a spare copy of every model. FailLite (Wu et al., 2025) makes model serving failure-resilient without that cost through heterogeneous replication. Instead of identical copies, it backs each model with cheaper, smaller variants, so a failed serving node degrades to a lighter model rather than dropping requests. This builds on an earlier result, that falling back to a smaller model through graceful service degradation (Hanafy et al., 2023) beats losing the request outright.

FailLite backs each model with smaller, cheaper variants, so a node failure degrades service to a lighter model instead of dropping requests.

The same principle extends beyond single models. It reaches the practical realities of deploying resilient ML at the edge (Practical Considerations (Gudipaty et al., 2025)) and full multi-stage inference pipelines rather than a single model (Resilient Distributed ML Inference Pipelines (Wu et al., 2024)).


Ensembles for Resilience

Heterogeneous backups help most when the models fail differently from one another. MEL (Gudipaty et al., 2025) trains a set of lightweight models that refine each other’s predictions when multiple servers are available, and each model still performs well on its own when failures cut the ensemble down to one. A diversity-encouraging training objective makes the models complementary rather than redundant, so the ensemble stays accurate as servers drop out.

MEL trains diverse lightweight models that collaboratively refine each other when servers are available and degrade gracefully to standalone models under failure.

Validating AI-Generated Decisions

Resilience keeps a system running. Safety keeps it from acting on bad decisions. As devices increasingly use large language models to configure and coordinate with unfamiliar peers on the fly, LLM output cannot be trusted blindly in a control loop.

CollabIoT (Shastri et al., 2025) makes LLM-driven configuration safe for transient IoT collaboration. It uses an LLM to translate a user’s high-level intent into fine-grained, capability-based access-control policies, validates those policies before they take effect, and confines LLM reasoning to a control plane while lightweight proxies enforce only validated decisions in a low-latency data plane.

This approach is part of our broader vision, the Internet of Collaborating Things (Hanafy et al., 2026), in which agentic edge AI negotiates and executes cross-domain collaborations while keeping unsafe or incorrect AI-generated decisions out of the control loop.

CollabIoT confines LLM policy generation and validation to a control plane, while lightweight proxies enforce only validated decisions in the data plane.

References

2026

  1. Under Review
    The Internet of Collaborating Things: Agentic Edge AI for Autonomous Cross-Domain Collaborations
    Walid A. Hanafy, Nader Sehatbakhsh, David Irwin, and 2 more authors
    Dec 2026

2025

  1. SoCC
    SoCC25.jpg
    FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
    Li Wu, Walid A. Hanafy, Tarek Abdelzaher, and 3 more authors
    In Proceedings of the 2025 ACM Symposium on Cloud Computing, Dec 2025
  2. MILCOM
    Practical Considerations for Failure Resilient ML Systems at the Edge
    Krishna Praneet Gudipaty, Walid A. Hanafy, Li Wu, and 6 more authors
    In MILCOM 2025 - 2025 IEEE Military Communications Conference (MILCOM), Dec 2025
  3. Arxiv
    MEL.jpg
    MEL: Multi-level Ensemble Learning for Resource-Constrained Environments
    Krishna Praneet Gudipaty, Walid A. Hanafy, Kaan Ozkara, and 4 more authors
    Dec 2025
  4. SEC
    SEC-LLM.jpg
    LLM-Driven Auto Configuration for Transient IoT Device Collaboration
    Hetvi Shastri, Walid A. Hanafy, Li Wu, and 3 more authors
    In Proceedings of the Tenth ACM/IEEE Symposium on Edge Computing (SEC), Dec 2025

2024

  1. MILCOM
    Enhancing Resilience in Distributed ML Inference Pipelines for Edge Computing
    Li Wu, Walid A. Hanafy, Abel Souza, and 3 more authors
    In MILCOM 2024 - 2024 IEEE Military Communications Conference (MILCOM), Dec 2024

2023

  1. MILCOM
    Failure-Resilient ML Inference at the Edge through Graceful Service Degradation
    Walid A. Hanafy, Li Wu, Tarek Abdelzaher, and 2 more authors
    In Proceedings of the 41st IEEE Military Communications Conference (MILCOM) workshop on Internet of Things for Adversarial Environments, Oct 2023