Fusion type multi-user audio and video communication method, system and device and storage medium
By selecting a high-performance client as the first node to forward and mix audio and video streams, the problem of device overload in existing technologies is solved, the bandwidth and CPU resource utilization of audio and video communication are improved, and the user experience is enhanced.
Patent Information
- Application Number
- CN202510989131.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-18
AI Technical Summary
In multi-person audio and video communication scenarios, existing Mesh, MCU and SFU architectures can easily overload servers or low-performance devices when building audio and video communication network connections, resulting in audio and video call stuttering and delays, which affects user experience.
Select clients that meet the set performance standards as the first node, and forward and mix audio and video streams through these nodes to avoid overloading the server and low-performance devices, and distribute audio and video processing tasks to various high-performance nodes.
It improves the bandwidth and CPU resource utilization of the audio and video communication architecture, and enhances the audio and video stream processing effect and user experience.
Smart Images

Figure CN120980183A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of communication, in particular to a fusion type multi-person audio and video communication method, system, device and storage medium. BACKGROUND
[0002] At present, in the multi-person audio and video communication scene, the Mesh (point-to-point direct connection) architecture, MCU (multi-point control unit) architecture or SFU (selective forwarding unit) architecture is usually used to build the audio and video communication network connection. Among them, in the Mesh architecture, each client establishes a direct P2P (peer-to-peer, peer-to-peer network) connection with all other clients, and its architecture is simple, which is suitable for small-scale audio and video conference scene; the MCU architecture receives all client audio and video streams through the center server, decodes, mixes and processes each audio and video stream, and then sends the synthesized single stream to each client. Its processing burden on the client is small, which is suitable for devices with limited client performance; the SFU architecture receives each client's audio and video stream through the server and forwards it to other clients as needed without decoding and mixing. Each client processes multiple audio and video streams, which consumes less server resources, has less delay, and supports flexible stream control and adaptive strategy. By adaptively selecting the above different audio and video communication structures, multi-person audio and video communication in the corresponding scene can be realized.
[0003] However, in the scene of multi-person audio and video communication in the live room, using the above Mesh architecture, MCU architecture or SFU architecture to build the audio and video communication network connection requires the server or the client to perform video stream decoding and mixing operation, which is easy to cause overload of the node with limited performance, affecting the audio and video stream processing effect, resulting in audio and video call lag, delay and other problems, and affecting the user's audio and video call experience. SUMMARY
[0004] Embodiments of the present application provide a fusion type multi-person audio and video communication method, system, device and storage medium, which selects a set number of clients meeting the set performance standard as first nodes, and the rest of the clients as second nodes, forwards the audio and video stream of the second node to the server through the first node, and mixes and forwards the audio and video stream issued by the server to the second node, so that the audio and video processing is dispersed to each first node, avoiding the overload of the server and the low-performance device, thereby improving the bandwidth and CPU resource utilization of the audio and video communication architecture, and solving the technical problem of node overload in the audio and video communication process. Compared with the existing audio and video communication architecture, the present application saves the bandwidth and CPU resources of the audio and video communication architecture, greatly improves the audio and video stream processing effect and the audio and video call experience.
[0005] In a first aspect, embodiments of the present application provide a fusion type multi-person audio and video communication method, comprising: obtaining specified device performance parameters of a plurality of clients currently initiating an audio and video call, selecting a set number of clients meeting a set performance standard as first nodes from the plurality of clients based on the specified device performance parameters, and selecting the remaining clients as second nodes, and controlling each first node to select a matched second node to build an audio and video communication connection; receiving first mixed audio and video streams uploaded by each first node, the first mixed audio and video streams being generated based on first audio and video streams of the first nodes and second audio and video streams uploaded by the corresponding second nodes to the first nodes; sending the first mixed audio and video streams uploaded by other first nodes to the current first node, so as to mix the received first mixed audio and video streams with second audio and video streams through the current first node to output second mixed audio and video streams, mix the received first mixed audio and video streams with first audio and video streams of the current first node to generate third mixed audio and video streams, and send the third mixed audio and video streams to the corresponding second nodes.
[0006] In a second aspect, embodiments of the present application provide a fusion type multi-person audio and video communication system, comprising: a node setting module configured to obtain specified device performance parameters of a plurality of clients currently initiating an audio and video call, select a set number of clients meeting a set performance standard as first nodes from the plurality of clients based on the specified device performance parameters, and select the remaining clients as second nodes, and control each first node to select a matched second node to build an audio and video communication connection; a receiving module configured to receive first mixed audio and video streams uploaded by each first node, the first mixed audio and video streams being generated based on first audio and video streams of the first nodes and second audio and video streams uploaded by the corresponding second nodes to the first nodes; a sending module configured to send the first mixed audio and video streams uploaded by other first nodes to the current first node, so as to mix the received first mixed audio and video streams with second audio and video streams through the current first node to output second mixed audio and video streams, mix the received first mixed audio and video streams with first audio and video streams of the current first node to generate third mixed audio and video streams, and send the third mixed audio and video streams to the corresponding second nodes.
[0007] In a third aspect, embodiments of the present application provide a fusion type multi-person audio and video communication device, comprising: a memory and one or more processors; the memory is configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the fusion type multi-person audio and video communication method according to the first aspect.
[0008] In a fourth aspect, the embodiments of the present application provide a non-volatile computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are configured to perform the fusion type multi-person audio and video communication method when executed by a computer processor.
[0009] In a fifth aspect, the embodiments of the present application provide a computer program product, which contains instructions, and when the instructions are run on a computer or processor, the computer or processor performs the fusion type multi-person audio and video communication method.
[0010] The embodiments of the present application select a set number of clients meeting the set performance standard as the first node and the rest of the clients as the second node based on the specified device performance parameters of the multiple clients currently initiating the audio and video call, control each first node to select a matched second node to build an audio and video communication connection, receive the first mixed audio and video stream uploaded by each first node, and the first mixed audio and video stream is generated based on the first audio and video stream of the first node and the second audio and video stream uploaded by the corresponding second node to the first node, send the first mixed audio and video stream uploaded by the other first node to the current first node, mix the received first mixed audio and video stream with the second audio and video stream through the current first node to output the second mixed audio and video stream, mix the received first mixed audio and video stream with the first audio and video stream of the self to generate the third mixed audio and video stream, and send the third mixed audio and video stream to the corresponding second node. By selecting a set number of clients meeting the set performance standard as the first node and the rest of the clients as the second node, forwarding the audio and video stream of the second node to the server by the first node, and mixing and forwarding the audio and video stream distributed by the server to the second node, the audio and video processing is dispersed to each first node, the overload of the server and the low-performance device is avoided, the bandwidth and CPU resource utilization of the audio and video communication architecture are improved, and the audio and video stream processing effect and the audio and video call experience are improved. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is a flowchart of a fusion type multi-person audio and video communication method provided by the embodiments of the present application; Figure 2 is a determination flowchart of the first node in the embodiments of the present application; Figure 3 is a connection schematic diagram of the server and the client in the embodiments of the present application; Figure 4 is an audio and video communication structure schematic diagram in the embodiments of the present application; Figure 5 is a structural schematic diagram of a fusion type multi-person audio and video communication system provided by an embodiment of the present application; Figure 6 is a structural schematic diagram of a fusion type multi-person audio and video communication device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0012] In order to make the objects, technical solutions and advantages of the present application clearer, the following further describes specific embodiments of the present application with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that, for the convenience of description, only parts related to the present application are shown in the drawings, but not all contents. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted by flowcharts. Although the flowcharts describe each operation (or step) as a sequential process, many of the operations can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0013] The fusion type multi-person audio and video communication method provided by the present application aims to select a set number of clients meeting the set performance standard as first nodes, and the rest of the clients as second nodes, forward the audio and video streams of the second nodes to the server through the first nodes, and mix and forward the audio and video streams issued by the server to the second nodes, so that the processing of the audio and video is dispersed to each first node, avoiding the overload of the server and low-performance devices.
[0014] In the multi-person audio and video communication scenario, the related multi-person audio and video communication methods are generally divided into three architectures: 1. Mesh (point-to-point direct connection) architecture, in which each participant (i.e. client) establishes a direct P2P connection with all other participants. Its advantage is that the architecture is simple, suitable for small-scale meetings, and the disadvantage is that as the number of participants increases, the number of connections increases exponentially, the bandwidth and CPU resource consumption rises rapidly, leading to a decrease in device performance, so the Mesh architecture is mainly suitable for small meeting scenarios with 2-3 people.
[0015] 2. MCU (Multipoint Control Unit) Architecture: The MCU architecture receives audio and video streams from all participants through a central server, decodes and mixes them, and then sends the synthesized single stream to each participant. Its advantages include low client-side processing load, making it suitable for devices with limited terminal performance. Its disadvantages include high resource consumption and higher latency due to the server's need for audio and video decoding and mixing. The MCU architecture is commonly used in traditional audio and video conferencing systems, but in scenarios with high real-time interaction requirements, its latency and resource consumption can become bottlenecks.
[0016] 3. SFU (Selective Forwarding Unit) architecture: In the SFU architecture, the server receives the audio and video streams from each participant and forwards them to other participants as needed, without decoding or mixing them. Its advantages include lower server resource consumption, lower latency, and support for flexible flow control and adaptive strategies. The disadvantage is that the client needs to process multiple audio and video streams, which may increase the terminal's workload.
[0017] In live streaming scenarios where multiple people speak on the microphone for audio and video calls, the Mesh architecture clearly requires establishing a large number of P2P connections, making it unsuitable for multi-person audio and video calls. The MCU architecture can handle scenarios with multiple people speaking on the microphone, but it consumes significant server resources. In SFU architecture scenarios with multiple people speaking on the microphone, the server needs to process and forward multiple streams. Assuming there are N people in a voice / video room, and all N people speak on the microphone, the server will receive N uplink streams and need to forward N*(N-1) streams, resulting in significant server resource consumption. Each client also needs to process N-1 streams, which is a major challenge for some devices with limited terminal performance.
[0018] For example, in a Mesh architecture, assuming there are N users speaking, each establishing a direct P2P connection with all other users, each user needs to send N-1 uplink streams and receive N-1 downlink streams. In an MCU architecture, assuming there are N users speaking, each connected to a central server, the central server receives all users' audio and video streams, decodes and mixes them, and then sends the synthesized single stream to each participant. The server needs to receive N uplink streams and send N downlink streams. Furthermore, the server needs to perform N mixing operations, which consumes significant machine resources. Each user needs to send one uplink stream and receive one downlink stream. In an SFU architecture, assuming there are N users speaking, each connected to a central server, the central server receives all users' audio and video streams but does not decode or mix them; instead, it forwards them to other participants. The server needs to receive N uplink streams and send N*(N-1) downlink streams. Each user needs to send one uplink stream and receive N-1 downlink streams.
[0019] Obviously, for the Mesh architecture, as the number of people on the microphone increases, the number of connections will increase exponentially, the bandwidth and CPU resource consumption will increase rapidly, leading to performance degradation, so the Mesh architecture is not suitable for multi-person audio and video call scenarios; for the MCU architecture, as the number of people on the microphone increases, the CPU resource consumption of the server will also increase rapidly, which requires dynamic expansion of more servers and increases the connection between servers; for the SFU architecture, as the server does not mix flow, it directly forwards all streams through the server, as the number of people on the microphone increases, the number of streams forwarded by the server will increase rapidly, and the number of streams received by the client will also increase accordingly, so the CPU resource consumption of the client for mixing flow will also increase rapidly, leading to performance degradation.
[0020] Therefore, for the above Mesh architecture, MCU architecture or SFU architecture, when building an audio and video communication network connection, the server or the client needs to perform video stream decoding and mixing operation, and for nodes with limited performance, overload may occur, affecting the audio and video stream processing effect, leading to audio and video call lag, delay and other problems, affecting the user's audio and video call experience. Based on this, the fusion multi-person audio and video communication method provided in the embodiments of the present application is provided to solve the technical problem of node overload in the audio and video communication process, so that the bandwidth and CPU resources of the entire system are better allocated and utilized, the bandwidth and CPU resource consumption of the entire system is reduced as a whole, and the overall user experience is improved.
[0021] Embodiment Figure 1 A flowchart of a fusion multi-person audio and video communication method provided in an embodiment of the present application is given. The fusion multi-person audio and video communication method provided in the embodiment can be executed by a fusion multi-person audio and video communication device. The fusion multi-person audio and video communication device can be realized by software and / or hardware. The fusion multi-person audio and video communication device can be composed of two or more physical entities, or can be composed of one physical entity. Generally, the fusion multi-person audio and video communication device can be a server device in a multi-person audio and video communication architecture.
[0022] The following takes a server as an example of the main body executing the fusion multi-person audio and video communication method for description. Referring to Figure 1 , the fusion multi-person audio and video communication method specifically includes: S110, obtaining specified device performance parameters of a plurality of clients currently initiating audio and video calls, selecting a set number of clients meeting a set performance standard as first nodes from the plurality of clients based on the specified device performance parameters, and selecting the remaining clients as second nodes, and controlling each first node to select a matched second node to build an audio and video communication connection.
[0023] When the application carries out multi-person audio and video communication, the server collects the index parameters of the identification device performance of all clients currently initiating the audio and video call to select the client with better performance to forward and mix the audio and video stream.
[0024] Among them, by collecting the specified device performance parameters of all clients in real time, the client with the best performance is selected according to the specified device performance parameters, and this part of the client is defined as the first node. The first node is used to directly connect with the server. For other clients, it is set as the second node. By controlling each first node to select the corresponding second node for connection, the audio and video communication connection is constructed. The uplink and downlink audio and video streams of the subsequent second node are all forwarded by the matched first node.
[0025] According to actual needs, the first node and the second node can be matched according to the proximity of the communication distance, or according to the actual network communication situation, a group of connections with the best signal can be selected for matching. In addition, in the case where the number of second nodes is more than that of first nodes, the first nodes with higher performance (i.e. the sorting of specified device performance parameters) can be configured with more second nodes, so as to ensure the reasonable use of node bandwidth and CPU resources.
[0026] Through the intelligent node selection mechanism, high-load tasks such as audio and video codec and stream mixing are accurately allocated to high-performance devices, effectively avoiding the lag phenomenon caused by the overload of low-performance terminals or central servers due to processing pressure, and reducing the network complexity by reducing the number of direct connections, forming a resource-optimized infrastructure, and laying a foundation for subsequent distributed processing.
[0027] Further, according to Figure 2 , the specified number of clients meeting the set performance standard are selected as the first node from the plurality of clients based on the specified device performance parameters, comprising: S1101, selecting the client meeting the set performance standard from the plurality of clients based on the specified device performance parameters, and sorting; S1102, selecting the specified number of clients from high to low as the first node from the clients meeting the set performance standard according to the sorting result of the specified device performance parameters.
[0028] By collecting the specified device performance parameters of all clients in real time, and then using a weighted scoring algorithm and other methods to perform multi-dimensional evaluation on the device, the candidate nodes that meet the set performance standards are screened out, such as clients with a weighted score higher than the set threshold. According to the actual specified device performance parameter settings, the corresponding performance standards can be adaptively set for candidate node screening. Then, according to the comprehensive performance score, the nodes are ranked in descending order, a priority queue is established by quantifying the device processing capacity, and it is ensured that only high-performance devices enter the candidate pool. The sorting mechanism provides differentiated basis for subsequent node allocation.
[0029] Then, according to the sorting result, the first node is dynamically selected from the first set number of clients, wherein the value of the set number can be determined by the ratio of the total load of the system to the processing limit of a single node, or can be determined according to the total number of clients. By preferentially selecting the most optimal performance device as the first node, a hybrid topology structure with high-performance terminals as the core is formed, so that the audio and video codec tasks are transferred to the first node for execution. It is ensured that the processing nodes on the critical path have sufficient redundant computing power, thereby reducing server bandwidth consumption in live streaming scenarios.
[0030] Optionally, the specified device performance parameters include at least one of CPU model, device memory, and network type. Screening clients that meet the set performance standards from a plurality of clients based on specified device performance parameters, comprising: Screening clients from a plurality of clients whose CPU model meets the set CPU model standard, device memory is greater than the set memory threshold, and / or network type meets the set model type.
[0031] In the node screening stage of the present application, the performance of the client is determined by multi-dimensional hardware features. Among them, based on the CPU model parameter (such as ARM Cortex-A78 and above architecture or equivalent performance x86 processor), the computing power baseline is established, combined with the device memory capacity threshold (such as setting at least 4GB available memory) and the network type whitelist (preferentially selecting terminals supporting 5GHz Wi-Fi6 or 5G SA network), a three-level filtering mechanism is constructed. The client that meets the conditions will enter the candidate pool as a candidate node. Then through a dynamic weighting algorithm, the CPU model, device memory and network type and other key indicators are quantitatively scored to form a continuous performance gradient sequence. By combining CPU model, device memory and network type and other indicator parameters to accurately analyze the performance of the client, accurate basic data is provided for node selection. In addition, according to actual needs, the specified device performance parameters can also be CPU computing power, network bandwidth and other parameters. In addition, when screening clients that meet the set performance standards, the client can be required to meet the corresponding performance standards of CPU model, device memory and network type, or only clients that meet one of the performance standards can be screened. The present application does not make fixed restrictions on the specific parameter types and client screening methods, and will not be described here.
[0032] Further, in the case that the total number of clients currently initiating audio and video calls is even, the set number is equal to one half of the total number of clients; Controlling each first node to select a matched second node to build an audio and video communication connection, comprising: Based on the peer-to-peer network, each first node is paired with a corresponding second node one-to-one to build an audio and video communication connection; In the case that the total number of clients currently initiating audio and video calls is odd, the set number is equal to one half of the total number of clients minus one; Controlling each first node to select a matched second node to build an audio and video communication connection, comprising: Based on the peer-to-peer network, each first node is paired with a corresponding second node one-to-one to build an audio and video communication connection; Correspondingly, after each first node is paired with a corresponding second node to build an audio and video communication, it further comprises: Setting the remaining unpaired second nodes as independent first nodes.
[0033] When screening the first node, if the number of clients N is even, the number of first nodes is set to 1 / 2 of the total number of clients through a dynamic topology adaptation mechanism, thereby forming a symmetrical processing architecture. Each high-performance first node establishes a P2P connection with a second node to build a bidirectional audio and video transmission link.
[0034] When the total number of clients is odd, the first node number is set to (N-1) / 2, and the complete pairing of N-1 clients is completed through the corresponding matching rules. The remaining unpaired second node is dynamically promoted to an independent first node.
[0035] As shown in Figure 3 , by screening the first node 12 and the second node 13, and then the first node 12 and the second node 13 are matched, the first node 12 is connected to the server 11, thereby constructing the audio and video call architecture.
[0036] Optionally, each first node is one-to-one paired with the corresponding second node based on the peer-to-peer network, including: The corresponding first node and the second node are selected based on the communication distance and the historical connection success rate, and the one-to-one pairing of the corresponding first node and the second node is performed based on the peer-to-peer network.
[0037] In the node pairing process of the present application, the communication quality between clients is dynamically evaluated through a network perception mechanism. First, the physical distance between nodes is calculated based on IP address resolution and geolocation technology, and the communication distance between the first node and the second node is accurately calculated by combining real-time network detection data (such as RTT delay, packet loss rate). At the same time, by analyzing the connection establishment success rate, interruption frequency and other stability indicators between the first node and the second node through historical connection records, a reliability scoring model is formed. Based on the above parameters, a weighted matching method is used to preferentially select the first node and the second node with excellent historical connection performance and adjacent communication distance for pairing. By minimizing the physical path length of data transmission, the end-to-end transmission delay is reduced, and by using historical behavior data to predict potential connection risks, the probability of audio and video stream interruption is reduced. In the live microphone connection scene, it can ensure that the node pair maintains a low two-way transmission delay, reduces the stuttering rate in a weak network environment, and significantly improves the stability and user experience of large-scale distributed audio and video communication.
[0038] Optionally, after obtaining the specified device performance parameters of the plurality of clients currently initiating the audio and video call, it further includes: If it is determined that the number of clients in the plurality of clients that meet the set performance standard is less than the set number based on the specified device performance parameters, the audio and video communication of the plurality of clients is constructed based on the set audio and video communication architecture.
[0039] For the scenario of insufficient node resources, the application also guarantees communication continuity through a dynamic architecture fallback mechanism. When it is detected that the number of clients meeting the preset performance standard is below a threshold, an architecture switching process is automatically triggered. For example, an SFU network is first attempted to be constructed, all original audio and video streams of the clients are received by the server and selectively forwarded, at this time the server only performs stream routing without mixing processing, and the clients need to handle the decoding and rendering of multiple streams by themselves; if the server load exceeds a safety threshold (such as CPU occupancy rate of 75%), the architecture is further degraded to an MCU architecture, and the central node completes all stream decoding, screen mixing and encoding operations to generate a single composite stream for distribution to each client; when the network condition deteriorates extremely, the architecture is finally downgraded to a Mesh architecture, and audio and video transmission is realized through fully connected P2P channels between the clients. Through the three-level fallback mechanism, the system can dynamically select the optimal degradation path according to the real-time resource state when the high-performance node is insufficient, ensuring that the audio and video session can still maintain the minimum availability under extreme conditions, thereby significantly enhancing the robustness in complex network environments.
[0040] S120, receiving the first mixed audio and video stream uploaded by each first node, the first mixed audio and video stream being generated based on the first audio and video stream of the first node and the second audio and video stream uploaded by the corresponding second node to the first node.
[0041] Further, based on the above audio and video communication architecture, during the audio and video call process, when the first node receives the original second audio and video stream uploaded by the second node managed by it through the P2P channel, it starts the mixed encoding process to perform space-time alignment and picture quality optimization on the first audio and video stream collected by itself and the second audio and video stream of the second node received.
[0042] Referring to Figure 4 , assuming that clients A, B, C and D perform a multi-person audio and video call, and A and D are first nodes respectively matching second nodes B and C. The audio and video stream of B is sent to A for processing, and the audio and video stream of A is mixed to generate an AB mixed audio and video stream, which is uploaded to the server 11 for forwarding. Similarly, the audio and video stream of C is sent to D for processing, and the audio and video stream of D is mixed to generate a CD mixed audio and video stream, which is uploaded to the server 11 for forwarding.
[0043] S130. Send the first mixed audio and video stream uploaded by other first nodes to the current first node, so that the received first mixed audio and video stream is mixed with the second audio and video stream by the current first node to output the second mixed audio and video stream, and the received first mixed audio and video stream is mixed with its own first audio and video stream to generate a third mixed audio and video stream, and the third mixed audio and video stream is sent to the corresponding second node.
[0044] After receiving the first mixed audio and video streams uploaded by all the first nodes, the server then forwards these streams to each first node. Understandably, a first node only needs to receive the first mixed audio and video streams from the other first nodes to obtain the audio and video stream information of all clients currently initiating the audio and video call. Therefore, when forwarding the first mixed audio and video streams, the server only needs to send the first mixed audio and video streams uploaded by the other first nodes, without needing to send the first mixed audio and video stream of the current first node itself.
[0045] Furthermore, the first node performs a secondary fusion of the received global first mixed audio and video stream with the original second audio and video stream uploaded by the second node it manages, thereby generating a second mixed audio and video stream containing the audio and video of all participants. This second mixed audio and video stream is then output to the local client. On the other hand, the first node overlays the received global first mixed audio and video stream with its own acquired first audio and video stream to generate a third mixed audio and video stream, which is then sent to the corresponding second node for output. This completes the audio and video stream processing and transmission of the entire audio and video communication architecture.
[0046] like Figure 4 As shown, node A receives a CD-mixed audio / video stream, then mixes it with its own audio / video stream to obtain an ACD-mixed audio / video stream, which is then forwarded to node B. Similarly, node D receives an AB-mixed audio / video stream, then mixes it with its own audio / video stream to obtain an ABD-mixed audio / video stream, which is then forwarded to node C. The entire encoding and mixing process is completed at the first node, avoiding the resource bottleneck of global mixing on the server side. Simultaneously, it ensures that the lower-performance client (the second node) only needs to decode a single mixed stream, significantly reducing the terminal processing pressure. Ultimately, through dynamic load balancing and a distributed processing architecture, the system's concurrency and overload resistance are improved while ensuring audio / video quality.
[0047] Exemplarily, based on the audio and video communication architecture of the present application, in the scene of multi-person microphone-up in a live room for audio and video call, assuming that there are N microphone-up users, first, select clients with better performance (such as high-end mobile phones, PC clients, etc.) as first nodes, these first nodes are directly connected with the server, and other clients are used as second nodes for P2P connection with the first nodes. By setting the matching rules, each first node is connected with at most one second node for P2P connection, so as to avoid too high load of the first node, and also allow those clients with limited performance to directly connect with the first node for P2P connection, which can also avoid too high load of the second node. Assuming that in the ideal case, half of the clients are first nodes directly connected with the server, and the other half are second nodes connected with the first nodes for P2P connection, that is, each first node is connected with one second node for P2P connection, the first node receives one stream sent by the second node, and then mixes the stream with the stream to be sent by itself before sending to the server, and the first node receives other streams from the server, mixes the streams with the stream to be sent by itself, and then sends to the second node. In the case of N microphone-up users (assuming that N is even), the server will receive N / 2 uplink streams, and then send N / 2 downlink streams, the first node will receive N / 2 streams, send 2 streams, mix 2 times, and the second node will receive 1 stream, send 1 stream, and not mix.
[0048] Based on this, by reasonably selecting half of the nodes as first nodes and the other half as second nodes, the mixing can be distributed to each first node. The N mixing processes are distributed to N / 2 first nodes, each first node mixes 2 times, which avoids the problem of concentrating the mixing pressure on the server, and each first node mixes at most 2 times, which also avoids too high load of the second node. At the same time, the number of received and sent streams of the server can be reduced, and the number of received streams of the client can be reduced, although the number of sent streams of the first node is increased by 1, but the total number of streams of the whole architecture is reduced, which can save the bandwidth and CPU resources of the server, save the bandwidth and CPU resources of the second node, fully utilize the bandwidth and CPU resources of the first node, make the bandwidth and CPU resources of the whole system better distributed and utilized, and reduce the bandwidth and CPU resource consumption of the whole system.
[0049] The above, by acquiring the specified device performance parameters of the multiple clients currently initiating the audio and video call, the specified number of clients meeting the set performance standard are selected as the first nodes from the multiple clients based on the specified device performance parameters, the remaining clients are the second nodes, each first node selects a matched second node to build an audio and video communication connection; receiving the first mixed audio and video stream uploaded by each first node, the first mixed audio and video stream is generated based on the first audio and video stream of the first node and the second audio and video stream uploaded by the corresponding second node to the first node; the first mixed audio and video stream uploaded by other first nodes is sent to the current first node, so as to mix the received first mixed audio and video stream and the second audio and video stream through the current first node to output the second mixed audio and video stream, mix the received first mixed audio and video stream and the first audio and video stream of itself to generate the third mixed audio and video stream, and send the third mixed audio and video stream to the corresponding second node. By using the above technical means, the set number of clients meeting the set performance standard are selected as the first nodes, the remaining clients are the second nodes, the audio and video stream of the second node is forwarded to the server by the first node, and the audio and video stream distributed by the server is mixed and forwarded to the second node, so that the processing of the audio and video is dispersed to each first node, the overload of the server and the low-performance device is avoided, the bandwidth and CPU resource utilization of the audio and video communication architecture are improved, and then the audio and video stream processing effect and the audio and video call experience are improved.
[0050] On the basis of the above embodiment, Figure 5 A structure schematic diagram of a fusion type multi-person audio and video communication system provided by the present application is provided. Referring to Figure 5 The fusion type multi-person audio and video communication system provided by the embodiment specifically includes: a node setting module 21, a receiving module 22 and a sending module 23.
[0051] The node setting module 21 is configured to acquire the specified device performance parameters of the multiple clients currently initiating the audio and video call, select the specified number of clients meeting the set performance standard as the first nodes from the multiple clients based on the specified device performance parameters, and control each first node to select a matched second node to build an audio and video communication connection; The receiving module 22 is configured to receive the first mixed audio and video stream uploaded by each first node, and the first mixed audio and video stream is generated based on the first audio and video stream of the first node and the second audio and video stream uploaded by the corresponding second node to the first node; The sending module 23 is configured to send the first mixed audio and video stream uploaded by other first nodes to the current first node, mix the received first mixed audio and video stream with the second audio and video stream to output a second mixed audio and video stream through the current first node, mix the received first mixed audio and video stream with the first audio and video stream of the current first node to generate a third mixed audio and video stream, and send the third mixed audio and video stream to the corresponding second node.
[0052] Specifically, a set number of clients meeting a set performance standard are selected as the first nodes from a plurality of clients based on a specified device performance parameter, including: The clients meeting the set performance standard are filtered from the plurality of clients based on the specified device performance parameter, and are sorted; The set number of clients are filtered from high to low among the clients meeting the set performance standard as the first nodes according to the sorting result of the specified device performance parameter.
[0053] The specified device performance parameter includes at least one of a CPU model, a device memory, and a network type; The clients meeting the set performance standard are filtered from the plurality of clients based on the specified device performance parameter, including: The clients whose CPU model meets a set CPU model standard, device memory is greater than a set memory threshold, and / or network type meets a set model type are filtered from the plurality of clients.
[0054] In a case where the total number of clients currently initiating an audio and video call is even, the set number is equal to one half of the total number of clients; The respective first nodes are controlled to select matched second nodes to build an audio and video communication connection, including: The respective first nodes are one-to-one paired with the corresponding second nodes based on a peer-to-peer network to build the audio and video communication connection; In a case where the total number of clients currently initiating an audio and video call is odd, the set number is equal to one half of the total number of clients minus one; The respective first nodes are controlled to select matched second nodes to build an audio and video communication connection, including: The respective first nodes are one-to-one paired with the corresponding second nodes based on a peer-to-peer network to build the audio and video communication connection; Correspondingly, after the respective first nodes are paired with the corresponding second nodes to build the audio and video communication, the method further includes: The remaining unpaired second nodes are set as independent first nodes.
[0055] The respective first nodes are one-to-one paired with the corresponding second nodes based on a peer-to-peer network, including: The corresponding first node and second node are selected based on the communication distance and the historical connection success rate, and the one-to-one pairing of the corresponding first node and second node is performed based on the peer-to-peer network.
[0056] In addition, after obtaining the specified device performance parameters of the multiple clients currently initiating the audio and video call, the method further comprises: Based on the specified device performance parameters, if the number of clients meeting the set performance standard is less than the set number, the audio and video communication of the multiple clients is constructed based on the set audio and video communication architecture.
[0057] The specified device performance parameters of the multiple clients currently initiating the audio and video call are obtained, the set number of clients meeting the set performance standard are selected as the first nodes from the multiple clients based on the specified device performance parameters, the remaining clients are selected as the second nodes, each first node selects a matched second node to construct an audio and video communication connection, the first mixed audio and video stream uploaded by each first node is received, the first mixed audio and video stream is generated based on the first audio and video stream of the first node and the second audio and video stream uploaded by the corresponding second node to the first node, the first mixed audio and video stream uploaded by the other first nodes is sent to the current first node, the received first mixed audio and video stream is mixed with the second audio and video stream to output a second mixed audio and video stream through the current first node, the received first mixed audio and video stream and the first audio and video stream of the current first node are mixed to generate a third mixed audio and video stream, and the third mixed audio and video stream is sent to the corresponding second node. By selecting the set number of clients meeting the set performance standard as the first nodes and the remaining clients as the second nodes, the audio and video stream of the second node is forwarded to the server by the first node, and the audio and video stream distributed by the server is mixed and forwarded to the second node, so that the processing of the audio and video is dispersed to each first node, the overload of the server and the low-performance device is avoided, the bandwidth and CPU resource utilization of the audio and video communication architecture are improved, and the audio and video stream processing effect and the audio and video call experience are improved.
[0058] The fusion-type multi-person audio and video communication system provided by the embodiments of the present application can be configured to perform the fusion-type multi-person audio and video communication method provided by the above embodiments, and has corresponding functions and advantages.
[0059] Based on the above actual examples, the embodiments of the present application further provide a fusion-type multi-person audio and video communication device, which is described with reference to Figure 6The fusion multi-person audio and video communication device includes a processor 31, a memory 32, a communication module 33, an input device 34, and an output device 35. The memory is a computer readable storage medium, and is configured to store software programs, computer executable programs, and modules, such as program instructions / modules of the fusion multi-person audio and video communication method according to any embodiment of the present application (for example, a node setting module, a receiving module, and a sending module in the fusion multi-person audio and video communication system). The communication module is configured to perform data transmission. The processor executes the software programs, instructions, and modules stored in the memory, thereby performing various function applications and data processing of the device, that is, implementing the fusion multi-person audio and video communication method described above. The input device is configured to receive input digital or character information, and to generate key signal input related to user settings and function control of the device. The output device can include a display device such as a display screen. The fusion multi-person audio and video communication device provided above can be configured to perform the fusion multi-person audio and video communication method provided in the above embodiments, and has corresponding functions and beneficial effects.
[0060] Based on the above embodiments, the present embodiment further provides a nonvolatile computer readable storage medium storing computer executable instructions, which are configured to perform a fusion multi-person audio and video communication method when executed by a computer processor. The storage medium can be any various types of memory device or storage device. Of course, the computer executable instructions of the nonvolatile computer readable storage medium provided in the present embodiment are not limited to the fusion multi-person audio and video communication method described above, but can also perform related operations in the fusion multi-person audio and video communication method provided in any embodiment of the present application.
[0061] Based on the above embodiments, the present embodiment further provides a computer program product. The technical solution of the present application, essentially or the part that contributes to the prior art, or the whole or part of the technical solution can be embodied in the form of a software product. The computer program product is stored in a storage medium, and includes a plurality of instructions for causing a computer device, a mobile terminal, or a processor therein to execute all or part of the steps of the fusion multi-person audio and video communication method described in the embodiments of the present application.
Claims
1. A method of fusion multi-person audio-video communication, characterized by, The method comprises: obtaining specified device performance parameters of a plurality of clients currently initiating an audio and video call, selecting a set number of clients meeting a set performance standard as first nodes from the plurality of clients based on the specified device performance parameters, and selecting the remaining clients as second nodes, and controlling each of the first nodes to select a matching second node to build an audio and video communication connection; receiving a first mixed audio and video stream uploaded by each of the first nodes, the first mixed audio and video stream being generated based on a first audio and video stream of the first node and a second audio and video stream uploaded by the corresponding second node to the first node; sending, to the current first node, the first mixed audio and video stream uploaded by the other first nodes, so as to mix the received first mixed audio and video stream with the second audio and video stream through the current first node to output a second mixed audio and video stream, mix the received first mixed audio and video stream with the first audio and video stream of the current first node to generate a third mixed audio and video stream, and send the third mixed audio and video stream to the corresponding second node.
2. The fusion multi-person audio-video communication method of claim 1, wherein, The method further comprises: selecting clients meeting a set performance standard from the plurality of clients based on the specified device performance parameters, and sorting the clients; selecting a set number of clients from high to low among the clients meeting the set performance standard according to the sorting result of the specified device performance parameters as the first nodes.
3. The fusion multi-person audio-video communication method of claim 2, wherein, The specified device performance parameters comprise at least one of a CPU model, device memory, and network type. The method further comprises: selecting clients whose CPU model meets a set CPU model standard, device memory is greater than a set memory threshold, and / or network type meets a set model type from the plurality of clients.
4. The fusion multi-person audio-video communication method of claim 1, wherein, In a case where the total number of clients currently initiating the audio and video call is even, the set number is equal to one half of the total number of clients. The method further comprises: pairing each of the first nodes with a corresponding second node one by one based on a peer-to-peer network to build the audio and video communication connection. In a case where the total number of clients currently initiating the audio and video call is odd, the set number is equal to one half of the total number of clients minus one. The method further comprises: pairing each of the first nodes with a corresponding second node one by one based on a peer-to-peer network to build the audio and video communication connection. Correspondingly, after the pairing, the method further comprises: setting the remaining unpaired second nodes as independent first nodes.
5. The fusion multi-person audio-video communication method of claim 4, wherein, The peer-to-peer network pairs each of the first nodes with a corresponding second node one-on-one, including: The corresponding first node and the second node are selected based on the communication distance and the historical connection success rate, and the peer-to-peer network is used to pair the corresponding first node and the second node one-on-one.
6. The fusion multi-person audio-video communication method according to any one of claims 1-5, characterized in that, After the specified device performance parameters of the multiple clients currently initiating the audio-video call are obtained, the method further includes: If the number of clients meeting the set performance standard is less than the set number based on the specified device performance parameters, the audio-video communication of the multiple clients is constructed based on the set audio-video communication architecture.
7. A converged multi-person audiovisual communication system, characterized by Including: A node setting module configured to obtain specified device performance parameters of multiple clients currently initiating an audio-video call, select a set number of clients meeting a set performance standard from the multiple clients as first nodes based on the specified device performance parameters, and select the remaining clients as second nodes, and control each of the first nodes to select a matching second node to construct an audio-video communication connection; A receiving module configured to receive a first mixed audio-video stream uploaded by each of the first nodes, the first mixed audio-video stream being generated based on a first audio-video stream of the first node and a second audio-video stream uploaded by the corresponding second node to the first node; A sending module configured to send the first mixed audio-video stream uploaded by other first nodes to the current first node, mix the received first mixed audio-video stream and the second audio-video stream to output a second mixed audio-video stream through the current first node, mix the received first mixed audio-video stream and the first audio-video stream of the current first node to generate a third mixed audio-video stream, and send the third mixed audio-video stream to the corresponding second node.
8. A converged multi-person audiovisual communication device, characterized by Including: A memory and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the fusion multi-person audio-video communication method according to any one of claims 1-6.
9. A non-transitory computer readable storage medium, comprising: The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by a computer processor, are configured to perform the fusion multi-person audio-video communication method according to any one of claims 1-6.
10. A computer program product, characterised in that, The computer program product contains instructions, which, when executed on a computer or processor, cause the computer or processor to perform the fusion multi-person audio-video communication method according to any one of claims 1-6.
Citation Information
Patent Citations
Multimedia conference realization method and device
CN102469409A
Method and system for processing conference call relay video
CN106559639A
Data transmission method and system, electronic device and computer readable storage medium
CN109495599A
Video conference implementation method and device, video conference system and storage medium
CN109660751A
Audio and video conference MCU cascading method, system and device and storage medium
CN113301292A