On July 29, 2026, beginning at approximately 9:07 AM EDT, Cloudli was notified of system alarms affecting portions of the Cloudli Connect voice platform. Two client hosts supporting registrar workloads experienced kernel panics caused by Media service modules operating in kernel mode. The incident affected SIP registrations and carrier connectivity and resulted in failed calls, including 603 “All Routes Exhausted” responses, as well as isolated calls with no audio issues.
Initial alarms were received at 9:07 AM EDT, and troubleshooting began at 9:10 AM. The unresponsive client host was rebooted at 9:15 AM. Cloudli posted investigating and identified updates at 9:21 AM and 9:23 AM, respectively. Test calls succeeded by 9:30 AM, and monitoring began at 9:33 AM, marking the system “Operational.” At this time Cloudli NOC continued monitoring the systems proactively as we began root cause analysis to implement a resolution. Cloudli issued a monitoring update at 12:47 PM notifying partners and customers of additional measures taken to mitigate the issue. The incident was formally resolved at 2:26 PM EDT.
Root cause analysis determined that client hosts experienced kernel panics associated with media service modules running in kernel mode for improved performance. The host failures disrupted registrar services and carrier routing-data access. Although cluster jobs were distributed across separate servers and datacenters, each registrar cluster relies on a single caching database instance, by design, for carrier database access. When the carrier cluster could not access the caching database, calls returned 603 “All Routes Exhausted” responses. The media service modules instability also caused intermittent audio failures.
Service was restored by rebooting the affected client hosts and restarting the media service module instances associated with audio failures. Engineering will continue investigating media service modules’ kernel-mode panic behavior and implement options that preserve performance and host stability. Cloudli will also complete additional redundancy work, including removal of the single-instance caching database dependency within each registrar cluster.
At Cloudli, we take any interruption of service very seriously and are continuously evaluating new processes and mitigation measures that can be proactively implemented to ensure service continuity.
When service interruptions do occur, our incident management procedure prioritizes prompt and clear notification, and timely status and resolution updates to our customers and partners.
We thank you for your continued support. Please feel free to reach out if you would like to discuss the particulars of this incident report further.