High Availability Deployment Architecture

Yeastar provides a High Availability (HA) solution for NovoOne deployment. Within the HA solution, NovoOne implements request distribution and synchronous processing to improve response speed, maintains normal operation of other nodes upon single-node failure to ensure system stability, and enables elastic scale-up of the system's service capacity based on business load.

Architecture

NovoOne adopts a cloud-native deployment architecture based on Kubernetes. This leverages container orchestration to support elastic workload scheduling, self-healing, and centralized resource management, supporting the stability and scalability of the overall call service.

The figure below shows the High Availability architecture of NovoOne.

Traffic source
  • Trunk Provider (ITSP): ITSP sends SIP signaling traffic and RTP media traffic for inbound and outbound PSTN calls.
  • SIP Endpoints: SIP endpoints include NovoOne Mobile Client and IP Phone. The endpoints send SIP signaling traffic and RTP media traffic for endpoint registration and call-related operations.
  • HTTPS Clients: HTTPS clients include NovoOne Platform, NovoOne Tenant, NovoOne Web Client, and Third-party platform integrated with NovoOne API. These clients send HTTPS traffic for web access, management operations, and API-based interactions.
Traffic loading balancing

All SIP signaling traffic and HTTPS traffic are routed to the Load Balancer Cluster, which has a fixed public IP address. This self-built cluster consists of two load balancer nodes (physical or virtual servers) - an active node and a standby node, providing high availability through failover.

Before forwarding ingress traffic to the Kubernetes Cluster, the active load balancer in the cluster runs health checks on backend services and distributes the traffic.

Traffic routing

Original RTP media traffic from traffic sources, together with SIP signaling traffic and HTTPS traffic forwarded by the Load Balancer Cluster, are routed to the Kubernetes Cluster and processed by corresponding components.

The cluster consists of four network-connected nodes (physical or virtual servers), with each node assigned a dedicated fixed public IP address, and all components run as Pods (the smallest deployable unit in Kubernetes) on these nodes.
Note: Double running Pods are deployed for each component type, delivering system high availability and fast responses to user requests. If one Pod of a component type fails, other running Pod instances of the same component type remain unaffected.

The cluster can schedule Pods to any node to meet the applicable scheduling requirements and enable traffic transmission between components hosted on different nodes.

Call limit

NovoOne adopting this deployment architecture supports up to 100 concurrent calls. To achieve higher call concurrency volumes, contact Yeastar Support for node scaling. For more information about architecture after node scaling, see Node scaling.

Node scaling

To achieve higher call concurrency volumes, scale out nodes to deploy additional Pod instances for certain component types for performance improvement. Examples of the newly deployed architecture are shown below.
Note: According to your business requirements, Yeastar Support scales out nodes, determines the quantity of Pods for each component type and which nodes the Pods land on.

For example, components such as FS-Media (media processing) and Monitor (system monitoring) consume substantial computing and memory resources; these components are deployed on dedicated nodes and excluded from cluster scheduling.