Cloudli Connect Service | Service Cloudli Connect

Incident Report for Cloudli Communications

Postmortem

Summary of Events:

On July 29, 2026, beginning at approximately 9:07 AM EDT, Cloudli was notified of system alarms affecting portions of the Cloudli Connect voice platform. Two client hosts supporting registrar workloads experienced kernel panics caused by Media service modules operating in kernel mode. The incident affected SIP registrations and carrier connectivity and resulted in failed calls, including 603 “All Routes Exhausted” responses, as well as isolated calls with no audio issues.

Incident Analysis and Mitigation Measures:

Initial alarms were received at 9:07 AM EDT, and troubleshooting began at 9:10 AM. The unresponsive client host was rebooted at 9:15 AM. Cloudli posted investigating and identified updates at 9:21 AM and 9:23 AM, respectively. Test calls succeeded by 9:30 AM, and monitoring began at 9:33 AM, marking the system “Operational.” At this time Cloudli NOC continued monitoring the systems proactively as we began root cause analysis to implement a resolution. Cloudli issued a monitoring update at 12:47 PM notifying partners and customers of additional measures taken to mitigate the issue. The incident was formally resolved at 2:26 PM EDT.

Root cause analysis determined that client hosts experienced kernel panics associated with media service modules running in kernel mode for improved performance. The host failures disrupted registrar services and carrier routing-data access. Although cluster jobs were distributed across separate servers and datacenters, each registrar cluster relies on a single caching database instance, by design, for carrier database access. When the carrier cluster could not access the caching database, calls returned 603 “All Routes Exhausted” responses. The media service modules instability also caused intermittent audio failures.

Service was restored by rebooting the affected client hosts and restarting the media service module instances associated with audio failures. Engineering will continue investigating media service modules’ kernel-mode panic behavior and implement options that preserve performance and host stability. Cloudli will also complete additional redundancy work, including removal of the single-instance caching database dependency within each registrar cluster.

Final Remarks:

At Cloudli, we take any interruption of service very seriously and are continuously evaluating new processes and mitigation measures that can be proactively implemented to ensure service continuity.

When service interruptions do occur, our incident management procedure prioritizes prompt and clear notification, and timely status and resolution updates to our customers and partners.

We thank you for your continued support. Please feel free to reach out if you would like to discuss the particulars of this incident report further.

Posted Jul 30, 2026 - 11:59 EDT

Resolved

This incident has been resolved. We will provide an RFO no later than three (3) business days from today.
***
Cet incident a été résolu. Nous fournirons un RFO au plus tard dans trois (3) jours ouvrables à partir d'aujourd'hui.
Posted Jul 29, 2026 - 14:26 EDT

Update

Additional improvements have been implemented, and we are continuing to monitor the results.
***
D'autres améliorations ont été apportées et nous continuons de surveiller les résultats.
Posted Jul 29, 2026 - 12:47 EDT

Monitoring

A fix has been implemented and we are monitoring the results.
***
Un correctif a été mis en œuvre et nous surveillons les résultats.
Posted Jul 29, 2026 - 09:33 EDT

Identified

The issue has been identified and fix is being implemented.

***

Le problème a été identifié et le correctif est en cours de mise en œuvre.
Posted Jul 29, 2026 - 09:23 EDT

Investigating

Dear clients,

We are currently investigating an issue with the Cloudli Connect service. This may impact some clients ability make and receive calls.

We will update you upon further discovery.

***

Chers clients,

Nous enquêtons actuellement sur un problème affectant le service Cloudli Connect. Cela pourrait affecter la capacité de certains clients à passer et recevoir des appels.

Nous vous tiendrons informés dès que nous aurons plus d'informations.
Posted Jul 29, 2026 - 09:21 EDT
This incident affected: Cloudli Connect (Inbound Calling, Outbound Calling) and Contact Centre (Other).