A method for reducing latency for a video group call

By adding negotiation and storage modules to the server and combining machine learning and reinforcement learning algorithms, the negotiation and resource configuration of video group calls are optimized, solving the latency problem of video group calls and improving user experience and communication efficiency.

CN122160474APending Publication Date: 2026-06-05SHANLITONGYI INFORMATION TECH (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANLITONGYI INFORMATION TECH (SHENZHEN) CO LTD
Filing Date
2026-03-27
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

The video call has a latency issue, resulting in a poor user experience and low communication efficiency, especially during frequent negotiation interactions and resource scheduling, which prolongs the communication effect.

Method used

By adding negotiation and negotiation storage modules to the server, pre-configuration and selective distributed nodes are implemented. Combined with machine learning and reinforcement learning algorithms, the negotiation process and resource allocation are optimized, reducing the interaction between the terminal and the server.

Benefits of technology

It significantly reduces video group call latency, improves immediacy and smoothness, simplifies the interaction between the terminal and the server, reduces server load, and enhances user experience and communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122160474A_ABST
    Figure CN122160474A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video group call, and particularly relates to a method for reducing video group call delay, which comprises the following steps: S1. creating a group; S2. group entering negotiation, when a user initially joins the group, a server initiatively initiates SDP negotiation and DTLS key negotiation with the user through a negotiation module, and saves negotiation results to a negotiation storage module; S3. initiating a session, any user in a session group initiates a video group call request to the server, the server judges the right of speech of the initiating user, after successful judgment, receives and analyzes the video stream of the initiating user according to the negotiation results stored in the negotiation storage module; and S4. video stream pulling, the server notifies other users to pull the stream, acquires the video stream and performs local playing. The application greatly reduces the interactive process of the terminal and the server, and effectively solves the delay problem of the video group call by combining the pre-configuration of the server and the selective establishment of a distributed node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video group call technology, and more specifically to a method for reducing video group call latency. Background Technology

[0002] Video group calling is a form of digital communication that has been widely used in many fields. Multiple terminal devices can join the same group through a network to conduct audio and video communication, providing people with a convenient and efficient way to communicate. For example, in remote conferencing scenarios, participants located in different geographical locations can join a conversation group using their terminal devices to achieve real-time video communication and collaboration.

[0003] Video group calls suffer from significant latency issues in practical applications, greatly impacting user experience and communication efficiency. Within a group, regardless of changes in video parameters and network configuration, each initiated video session requires negotiation with the server, leading to frequent communication interactions between the terminal and server. This results in prolonged waiting times for video group calls, increasing communication overhead. Furthermore, due to the complexity of these operations, each time a video group call is initiated, the server and terminal devices need to handle a large number of requests and responses, further burdening the system and increasing latency. This can even lead to audio-visual desynchronization, video stuttering, and other problems, resulting in untimely information delivery and severely impacting communication effectiveness and work efficiency. Summary of the Invention

[0004] To address the severe latency issue in video group calls, this invention provides a method for reducing video group call latency. This method significantly reduces the interaction process between the terminal and the server, and by combining server pre-configuration and selective establishment of distributed nodes, it significantly improves the immediacy and smoothness of video group calls initiated by users through the terminal, effectively solving the latency problem of video group calls.

[0005] The technical solution adopted in this invention is to provide a method for reducing video group call latency, including a session system composed of multiple terminals and a server. The server includes a negotiation module for negotiating video parameters and a negotiation storage module for storing negotiation data. Based on this, the method includes the following steps: S1. Create a group: The user sends a group creation request to the server through the terminal and sets the basic information and permission information of the group. After the server verifies the request successfully, a conversation group with a unique identifier and supporting audio and video is generated. S2. Group entry and negotiation: When a user joins a group for the first time, the server actively initiates SDP negotiation and DTLS key negotiation with the user through the negotiation module, and saves the negotiation results to the negotiation storage module. S3. Initiate a session: Any user in the session group initiates a video group call request to the server. The server determines the right to speak for the initiating user. If the determination is successful, the server receives and parses the video stream of the initiating user based on the negotiation results stored in the negotiation storage module. S4. Video streaming: The server notifies other users to stream the video. Other users obtain the video stream and play it locally based on the negotiation results stored in the negotiation storage module.

[0006] It also includes adding a negotiation triggering unit to the negotiation module. When the parameters of SDP negotiation or DTLS key negotiation change or need to be changed, the user can actively initiate SDP negotiation or DTLS key negotiation to the server by operating the negotiation triggering unit.

[0007] Step S2 also includes that when a user joins a group for the first time, the terminal actively reports its hardware configuration, network configuration, and user session history data to the server. Based on machine learning algorithms, the system predicts the user's resource requirements and network change trends in video group calls, and pre-configures network resources for the user based on the prediction results.

[0008] Step S2 also includes the server evaluating the terminal performance based on the user's terminal hardware configuration and network configuration, selecting a terminal that meets the conditions as a distributed processing node, which shares the processing results and intermediate data with the server.

[0009] Step S3 also includes introducing a reinforcement learning algorithm during the video group call process, where the server dynamically adjusts the network resource configuration of each terminal based on real-time data from each terminal and distributed processing node.

[0010] The beneficial effects of this invention are that it provides a method to reduce video group call latency. When a user joins a group for the first time through a terminal, the server actively negotiates with the user's terminal and stores the negotiation result as the negotiation parameters for the user's video group call. This eliminates the need for negotiation every time a session is initiated. Furthermore, by combining server pre-configuration and dynamic adjustment, as well as selectively establishing distributed nodes on some terminals, the immediacy and smoothness of video group calls are significantly improved, effectively solving the latency problem of existing video group calls. Attached Figure Description

[0011] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0012] like Figure 1As shown, this invention provides a method for reducing video group call latency, including a session system composed of multiple terminals and a server. The server includes a negotiation module for negotiating video parameters and a negotiation storage module for storing negotiation data. A negotiation triggering unit is added to the negotiation module. When negotiation parameters change or need to be changed, the user can actively initiate SDP negotiation or DTLS key negotiation with the server by operating the negotiation triggering unit through the terminal. Based on this, the method includes the following steps: S1. Create a group Users initiate a group creation request to the server via their terminals, setting basic group information and permissions, including group name, group type, group joining rules, and group member permissions. Upon receiving the group creation request, including user information, basic group information, and permission information, the server verifies the user information and basic group information. After successful verification, a uniquely identified session group supporting audio and video is generated. The server saves the session group record to the database and generates a unique index for association, which is used for later querying of session group records. At the same time, the server generates a configuration file for the session group to store various group settings and status information.

[0013] After the server completes the creation of the session group, it generates a response data packet containing creation success information, group identifier, group type, basic group information, and permission information and sends it to the terminal.

[0014] S2. Group Entry and Negotiation Other users join the session group through their respective terminals. When the negotiation module detects a new member joining, it proactively initiates an SDP negotiation request to the new member. This request contains information such as the audio and video encoding formats supported by the server, parameter ranges, and transmission protocols, which is encapsulated in JSON format and sent to the terminal. For example, the server supports H.264 video encoding format, Opus audio encoding format, video resolution range of [640×480, 1920×1080], and frame rate range of [15fps, 60fps], etc.

[0015] Upon receiving the SDP negotiation request, the terminal parses the request content and, based on its own hardware and software capabilities, selects the supported audio and video encoding formats, parameters, and transmission protocols. For example, the terminal device supports H.264 video encoding, Opus audio encoding, and supports video resolutions of [1280×720, 1920×1080], frame rates of [30fps, 60fps], and support transmission protocols of RTP / RTCP and QUIC. Then, it constructs a response data packet, feeding back the terminal's supported information to the negotiation module.

[0016] After receiving the terminal's response, the negotiation module analyzes the commonalities supported by both parties and prioritizes the option that satisfies the capabilities of most terminals while ensuring good audio and video quality and transmission efficiency. For example, regarding video encoding format, H.264 is chosen because of its better compatibility and support by both parties; for audio encoding, Opus is selected considering sound quality and bandwidth requirements; the video resolution is chosen based on the overall group situation and network conditions; and the frame rate is selected at 30fps to ensure smooth video playback without consuming excessive bandwidth. Regarding the transmission protocol, given the complexity of the current network environment and the requirements for transmission efficiency, the QUIC protocol is selected.

[0017] The negotiation module encapsulates the final negotiated audio and video encoding formats, parameters, and transmission protocols into an acknowledgment data packet and sends it to the terminal device. After receiving the acknowledgment data packet, the terminal device parses and confirms that the negotiation result is consistent with the local configuration. After confirmation, it sends an acknowledgment message to the negotiation module, indicating that the SDP negotiation is complete.

[0018] After the SDP negotiation is completed, the negotiation module initiates the DTLS key negotiation process with the new member. The negotiation module sends a DTLSHello message to the terminal device, which includes the server's random number ServerRandom, a list of supported cipher suites such as AES-256-GCM and ChaCha20-Poly1305, and other relevant parameters.

[0019] After receiving the Hello message, the terminal device selects a supported cipher suite from the list of cipher suites provided by the negotiation module, and generates its own random number, ClientRandom. Then, the terminal device sends a Hello message containing ClientRandom and the selected cipher suite to the negotiation module as a response.

[0020] After receiving the terminal's response, the negotiation module confirms that the encryption suite selected by the terminal is within its supported range, and ultimately determines the encryption suite to be used from the encryption suites supported by both parties. It then sends its own digital certificate to the terminal, which contains the server's public key and related authentication information, used by the terminal device to verify the server's identity.

[0021] After receiving the server's certificate, the terminal uses its built-in certificate verification mechanism to verify the server certificate's legitimacy through a trusted root certificate. Once the verification is successful, the terminal device generates a Pre-Master Secret, encrypts the Pre-Master Secret using the public key from the server certificate, and then sends the encrypted Pre-Master Secret to the server via a key exchange message.

[0022] After receiving the encrypted pre-master key from the terminal, the negotiation module decrypts it using its own private key to obtain the pre-master key. Then, the server and the terminal device, based on the pre-master key and the previously exchanged ServerRandom and ClientRandom, respectively, generate a session key (SessionKey) using a specific key derivation function, such as HKDF-SHA256. This session key will be used for the encryption and decryption of subsequent communication data within this session group.

[0023] Both the server and the terminal use their generated session keys to perform encrypted communication tests to ensure the keys are valid. Upon successful testing, both parties send a handshake completion message to each other, completing the DTLS key negotiation. At this point, a secure communication channel is established between the server and the new member terminal. Subsequent data transmission during video group calls will be conducted through this encrypted channel, ensuring data confidentiality and integrity.

[0024] After completing SDP and DTLS key negotiation, the terminal proactively reports its hardware configuration, network configuration, and user session history data to the server, including CPU model, number of cores and computing power, GPU model and graphics processing power, maximum supported texture resolution, storage capacity and read / write speed, and network interface type and bandwidth capacity. Based on machine learning algorithms and terminal performance, the server predicts the user's resource requirements and network change trends during video group calls, and pre-configures personalized resources for the user based on the prediction results.

[0025] Personalized resources include network resources and storage resources. Network resources include bandwidth resources and network connectivity resources, used for data interaction between terminals and servers. Storage resources include cache space and storage quotas.

[0026] After resource pre-configuration, the server evaluates terminal performance based on the user's terminal hardware and network configurations. Based on preset node conditions, including computing power, storage capacity, network conditions, and geographical location, the server selects terminals that meet these conditions and sends them invitation messages to become distributed nodes. This includes the responsibilities and resource consumption of becoming a distributed node. Upon receiving the invitation, the terminal negotiates with the server regarding task allocation, resource usage, and data interaction methods. For example, the terminal negotiates with the server the types of tasks it can handle, such as video encoding and edge caching, as well as the additional resources required during task execution, such as increased bandwidth. If the negotiation is successful, the terminal confirms its status as a distributed node; as a distributed processing node, the terminal shares processing results and intermediate data with the server.

[0027] During the video group call process, the server introduces a reinforcement learning algorithm to dynamically adjust the resource configuration of each terminal based on real-time data from each terminal and distributed processing node.

[0028] S3. Initiate a session: Any user in the session group initiates a video group call request to the server. The server determines the right to speak for the initiating user and checks whether any other member currently holds the right to speak. If no member holds the right to speak, the right to speak is allocated to the user who initiated the request. If a member holds the right to speak, the server determines whether to preempt the right to speak based on other rules.

[0029] Once the user who initiated the request has obtained the right to speak, the server receives and parses the video stream from the user based on the negotiation results stored in the negotiation storage module. S4. Video streaming: The server notifies other users to stream the video. Other users obtain the video stream and play it locally based on the negotiation results stored in the negotiation storage module.

[0030] This invention first adds a negotiation module and a negotiation storage module to the server. After a group is created, when a user joins the group via a terminal, the server actively negotiates with the newly joined user's terminal through the negotiation module, including SDP negotiation and DTLS key negotiation. The negotiation results are stored in the negotiation storage module as interaction parameters between the terminal and the server when the user subsequently initiates a video group call. Therefore, when the user subsequently initiates a video group call within the group, no further negotiation is required; the negotiation results stored in the negotiation storage module can be used directly. This significantly simplifies the interaction process between the terminal and the server during video group calls and effectively reduces the latency of initiating video group calls.

[0031] The present invention also adds a negotiation triggering module to the server. When the user needs to change the negotiation parameters, the terminal can trigger the negotiation triggering module to re-initiate the SDP negotiation or DTLS key negotiation with the server, which improves the convenience and flexibility of the interaction between the terminal and the server.

[0032] This invention not only negotiates when a user joins a group, but also pre-configures personalized resources for the terminal based on its performance after negotiation. This eliminates the need for the server to perform resource scheduling every time a video group call is initiated, further reducing latency when initiating a video group call and also reducing server load.

[0033] This invention evaluates terminal performance during pre-configuration and selects terminals that meet set conditions as distributed nodes to share results and intermediate data with the server. The server can assign some tasks to these distributed nodes for processing, thereby improving the server's data processing speed. This further increases the data processing speed when users in a group initiate video group calls. Combined with the pre-negotiation and pre-configuration method, it further reduces the latency when initiating video group calls.

[0034] This invention introduces a reinforcement learning algorithm into the video group call process, dynamically adjusting the configuration resources based on real-time monitored video group call data, thereby further ensuring the smooth flow of the video group call process.

Claims

1. A method for reducing video group call latency, comprising a session system consisting of multiple terminals and a server, characterized in that: The method includes the following steps: A negotiation module for negotiating video parameters and a negotiation storage module for storing negotiated data are added to the server. S1. Create a group: The user sends a group creation request to the server through the terminal and sets the basic information and permission information of the group. After the server verifies the request successfully, a conversation group with a unique identifier and supporting audio and video is generated. S2. Group entry and negotiation: When a user joins a group for the first time, the server actively initiates SDP negotiation and DTLS key negotiation with the user through the negotiation module, and saves the negotiation results to the negotiation storage module. S3. Initiate a session: Any user in the session group initiates a video group call request to the server. The server determines the right to speak for the initiating user. If the determination is successful, the server receives and parses the video stream of the initiating user based on the negotiation results stored in the negotiation storage module. S4. Video streaming: The server notifies other users to stream the video. Other users obtain the video stream and play it locally based on the negotiation results stored in the negotiation storage module.

2. The method for reducing video group call latency according to claim 1, characterized in that: It also includes adding a negotiation triggering unit to the negotiation module. When the parameters of SDP negotiation or DTLS key negotiation change or need to be changed, the user can actively initiate SDP negotiation or DTLS key negotiation to the server by operating the negotiation triggering unit.

3. The method for reducing video group call latency according to claim 1, characterized in that: Step S2 also includes that when a user joins a group for the first time, the terminal actively reports its hardware configuration, network configuration, and user session history data to the server. Based on machine learning algorithms, the server predicts the user's resource requirements and network change trends in video group calls, and pre-configures network resources for the user based on the prediction results.

4. The method for reducing video group call latency according to claim 2, characterized in that: Step S2 also includes the server evaluating the terminal performance based on the user's terminal hardware configuration and network configuration, selecting a terminal that meets the conditions as a distributed processing node, which shares the processing results and intermediate data with the server.

5. The method for reducing video group call latency according to claim 3, characterized in that: Step S3 also includes introducing a reinforcement learning algorithm during the video group call process, where the server dynamically adjusts the network resource configuration of each terminal based on real-time data from each terminal and distributed processing node.