Audio mixing methods, devices, edge servers, and central servers
By dynamically adjusting the audio mixing strategy based on client business information, the problems of high traffic overhead and long latency in streaming media forwarding are solved, achieving low-latency and efficient audio interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE M2M
- Filing Date
- 2022-03-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies for streaming media forwarding suffer from high traffic overhead and long latency, especially when a large number of forwards are performed within a streaming media cluster, resulting in high traffic overhead and cluster costs, while the cumulative latency gradually increases.
By obtaining the client's business information, the audio mixing strategy is determined to be either an edge-convergence mixing strategy, a multi-edge mixing center-convergence strategy, or a full-transmission forwarding strategy. The mixing strategy is dynamically adjusted according to the business scenario and server load to optimize the audio stream processing and reduce unnecessary data transmission.
It effectively reduces traffic overhead and latency during audio interaction, optimizes server load capacity and cost, and meets the low latency requirements of different business scenarios.
Smart Images

Figure CN116827911B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio mixing method, apparatus, edge server, and central server. Background Technology
[0002] Traditional streaming media aggregation solutions are implemented using MCUs and are widely used in video conferencing. Among the mainstream open-source streaming media servers based on the Web Real-time Communication (WebRTC) protocol, there are SFU servers, represented by international providers like JANUS and MediaSoup, and domestic open-source server SRS, all of which can solve the problem of streaming media aggregation and distribution in the cloud. With business growth, single-point services can hardly handle a large number of application access requests. To improve server capacity, cross-regional capabilities, and cross-network capabilities, the evolution naturally leads to streaming media service clusters that provide streaming media aggregation and forwarding in the cloud. While streaming media service clusters forward all incoming streaming media, offering good interactive features, large-scale distribution requires extensive forwarding within the cluster, resulting in high bandwidth consumption and high cluster costs. Furthermore, the cumulative latency gradually increases with each relay on the server. Summary of the Invention
[0003] The purpose of this application is to provide an audio mixing method, apparatus, edge server, and central server, thereby solving the problems of high traffic overhead and long latency in the prior art for streaming media forwarding.
[0004] Firstly, in order to achieve the above objectives, embodiments of this application provide an audio mixing method applied to a first edge server, comprising:
[0005] Obtain the client's business information;
[0006] Determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge-converging center-converging strategy, and full-forwarding strategy.
[0007] Optionally, determining the audio mixing strategy corresponding to the service information includes:
[0008] Based on the business scenario determined by the aforementioned business information, the audio mixing strategy is determined; and / or
[0009] When a preset proportion of business participants are connected to the same edge server based on the business information, the audio mixing strategy is determined to be an edge convergence mixing strategy.
[0010] Optionally, determining the audio mixing strategy based on the business scenario identified by the business information includes any one of the following:
[0011] When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy;
[0012] When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy.
[0013] When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
[0014] Optionally, the method further includes:
[0015] When the audio mixing strategy is an edge convergence mixing strategy, if the edge mixing function of the first edge server is enabled, the first input audio stream and the second input audio stream received from at least one second edge server are mixed to obtain a full mixed stream.
[0016] The full mixed stream is filtered to obtain the output stream corresponding to each client connected to the first edge server;
[0017] The full mixed stream is sent to each of the second edge servers and the output stream is sent to the corresponding client.
[0018] Wherein, the first input audio stream is the audio stream of the client connected to the first edge server, and the second edge server is a server whose edge mixing function is not enabled.
[0019] Optionally, the method further includes:
[0020] Obtain the carrying capacity of each of the second edge servers;
[0021] Based on the CPU load of the first edge server and its carrying capacity, the audio mixing strategy is controlled to switch from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy.
[0022] Optionally, controlling the audio mixing strategy to switch from the edge-converging mixing strategy to the multi-edge mixing center-converging strategy based on the CPU load of the first edge server and the carrying capacity includes:
[0023] When it is determined that the CPU load has reached a first preset load value, and / or when it is sensed that the number of clients connected to at least one of the second edge servers has reached a first preset number of connections, the audio mixing strategy is controlled to smoothly transition from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy.
[0024] Optionally, controlling the audio mixing strategy to smoothly transition from the edge-converging mixing strategy to the multi-edge-converging center-converging strategy includes:
[0025] The message queue server RabbitMQ is used to request the central server to start and notify the second edge server to start the edge hybrid function.
[0026] Optionally, the method further includes:
[0027] When the audio mixing strategy is an edge convergence mixing strategy, if the first edge server does not enable the edge mixing function, the first input audio stream is sent to the third edge server so that the third edge server can perform audio mixing.
[0028] Wherein, the first input audio stream is the audio stream of the client connected to the first edge server, and the third edge server is a server with edge mixing function enabled.
[0029] Optionally, after the step of sending the first input audio stream to the third edge server, the method further includes:
[0030] Receive the full mixed stream sent by the third edge server;
[0031] The full mixed stream is filtered to obtain the output stream corresponding to each client connected to the first edge server;
[0032] The output stream is sent to the client corresponding to the output stream.
[0033] Optionally, the method further includes:
[0034] When the audio mixing strategy is a multi-edge mixing center convergence strategy, the first input audio streams of each client connected to the first edge server are mixed to obtain a first mixed stream;
[0035] Send a mixing request to the central server, the mixing request carrying the first mixed stream;
[0036] The central server receives a full mixed stream, wherein the full mixed stream is generated by the central server by mixing the received first mixed stream with a second mixed stream sent by other edge servers.
[0037] The full mixed stream is filtered to obtain the output stream corresponding to the client connected to the first edge server;
[0038] The output stream is sent to the client corresponding to the output stream.
[0039] Optionally, the method further includes:
[0040] The first request message sent by the central server is received, which is used to instruct the first edge server to keep the edge blending function enabled and to disable the edge blending function of other edge servers.
[0041] In response to the first request message, the edge blending function is kept on to smoothly transition from the multi-edge blending center convergence strategy to the edge convergence blending strategy;
[0042] Wherein, the first request message is a message sent by the central server when it senses that the number of clients in the first channel has dropped to a third preset access number; the first channel is the channel where the clients accessing the first edge server are located; and the number of clients accessing the first edge server is the largest.
[0043] Optionally, a smooth transition from the multi-edge blending center convergence strategy to the edge convergence blending strategy includes:
[0044] Receive a fourth input audio stream sent by other edge servers within the first channel;
[0045] The received fourth input audio stream and the first input audio stream are mixed;
[0046] Send the mixed stream to the other edge servers and stop sending the audio stream to the central server.
[0047] Optionally, the method further includes:
[0048] The server receives a second request from the central server, which instructs the fourth edge server to keep the edge blending function enabled and disable the edge blending function of other edge servers.
[0049] In response to the second request, the edge blending function is disabled;
[0050] Receive a third request message from the fourth edge server requesting the first input audio stream;
[0051] In response to the third request message, the first input audio stream is sent to the fourth edge server;
[0052] Receive the mixed stream sent by the fourth edge server and stop sending the first input audio stream to the central server;
[0053] The fourth edge server is the server with the largest number of clients accessing the first channel, and the first channel is the channel where the clients accessing the first edge server are located.
[0054] Optionally, the full mixed stream is filtered to obtain the output stream corresponding to each client connected to the first edge server, including:
[0055] Based on the demand information of each client connected to the first edge server, a mute queue corresponding to each client is generated, and the mute queue is used to indicate audio streams that the client does not need;
[0056] Based on the mute queue, the audio streams related to the mute queue in the full mix stream are filtered out to obtain the output stream.
[0057] Optionally, the method further includes:
[0058] When the first edge server is an edge server and the audio mixing strategy is a full forwarding strategy, a third input audio stream from a preset client is requested from the fifth edge server;
[0059] The received third input audio is sent to each client connected to the first edge server.
[0060] Secondly, in order to achieve the above objectives, embodiments of this application also provide an audio mixing method applied to a central server, comprising:
[0061] Obtain the client's business information;
[0062] Determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge-converging center-converging strategy, and full-forwarding strategy.
[0063] Optionally, determining the audio mixing strategy corresponding to the service information includes:
[0064] Based on the business scenario determined by the aforementioned business information, the audio mixing strategy is determined; and / or
[0065] When a preset proportion of business participants are connected to the same edge server based on the business information, the audio mixing strategy is determined to be an edge convergence mixing strategy.
[0066] Optionally, determining the audio mixing strategy based on the business scenario identified by the business information includes any one of the following:
[0067] When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy;
[0068] When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy.
[0069] When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
[0070] Optionally, the method further includes:
[0071] When the audio mixing strategy is a multi-edge mixing center convergence strategy, the third mixed stream sent by each edge server is obtained;
[0072] The various third mixing streams are mixed to obtain the full mixing stream;
[0073] The full mixed stream is sent to each of the edge servers.
[0074] Optionally, the method further includes:
[0075] When the audio mixing strategy is an edge-convergence mixing strategy, a request message to enable the central convergence function of the central server is received from the fifth edge server.
[0076] In response to the request message, the central aggregation function is enabled;
[0077] When the central aggregation function is enabled, it receives the mixed stream sent by each of the fifth edge servers;
[0078] The full mixed stream obtained by mixing the mixed streams sent by each of the fifth edge servers is sent to each of the edge servers.
[0079] Thirdly, in order to achieve the above objectives, embodiments of this application also provide an audio mixing apparatus applied to a first edge server, comprising:
[0080] The first acquisition module is used to acquire the client's business information;
[0081] The determination module is used to determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge mixing center-converging strategy, and full forwarding strategy.
[0082] Fourthly, in order to achieve the above objectives, embodiments of this application also provide an audio mixing device applied to a central server, comprising:
[0083] The first acquisition module is used to acquire the client's business information;
[0084] The determination module is used to determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge-converging center-converging strategy, and full-forwarding strategy.
[0085] Fifthly, in order to achieve the above objectives, embodiments of this application also provide an edge server, including: a transceiver, a processor, a memory, and a program or instructions stored in the memory and executable on the processor; when the processor executes the program or instructions, it implements the steps of the audio mixing method as described in the first aspect.
[0086] Sixthly, in order to achieve the above objectives, embodiments of this application also provide a central server, including: a transceiver, a processor, a memory, and a program or instructions stored in the memory and executable on the processor; when the processor executes the program or instructions, it implements the steps of the audio mixing method as described in the second aspect.
[0087] Seventhly, in order to achieve the above objectives, embodiments of this application also provide a readable storage medium storing a program that, when executed by a processor, implements the steps of the audio mixing method as described in the first aspect, or the steps of the audio mixing method as described in the second aspect.
[0088] The above-mentioned technical solution of this application has at least the following beneficial effects:
[0089] This application provides an audio mixing method applied to a first edge server. The method includes: obtaining the client's business information and determining the audio mixing strategy corresponding to the business information. In this way, different audio mixing strategies are used to mix audio for different business information. Subsequently, the mixed audio is forwarded for interaction. This solves the problems of high traffic overhead and high latency that exist when using full audio forwarding for interaction in the prior art. Attached Figure Description
[0090] Figure 1 This is a flowchart illustrating one embodiment of the audio mixing method of this application;
[0091] Figure 2 This is a schematic diagram of the server cluster structure in the embodiments of this application;
[0092] Figure 3 This is a second schematic flowchart of an embodiment of the audio mixing method of this application;
[0093] Figure 4 This is the third flowchart illustrating an embodiment of the audio mixing method of this application.
[0094] Figure 5 This is a schematic diagram of one embodiment of the audio mixing device according to this application;
[0095] Figure 6 This is a second schematic diagram of the structure of an audio mixing device embodiment of this application;
[0096] Figure 7 This is a schematic diagram of the edge server structure according to an embodiment of this application. Detailed Implementation
[0097] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0098] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0099] Before providing a detailed description of the embodiments of this application, the relevant technologies will be explained first:
[0100] I. With the development of the times and the large-scale popularization of 5G, audio and video are moving towards a low-latency era. A large number of low-latency interactive audio and video applications have emerged. The demand for specific scenarios such as conferencing, live video conferencing, and remote education is increasing daily. WebRTC is an end-to-end low-latency audio and video transmission protocol stack open-sourced by Google, composed of a series of RFC standards. Due to its low-latency characteristics and widespread browser support, WebRTC-based streaming media applications dominate real-time interactive scenarios. Since it only implements end-to-end interaction, practical applications require the addition of a WebRTC streaming media server to better realize cloud functions such as media forwarding, protocol conversion, storage recording, and audio and video processing to meet various application scenarios. In summary, it is necessary to aggregate or distribute client-side streaming media in the cloud.
[0101] Traditional media aggregation solutions, represented by streaming media servers implemented using Multipoint Conferencing Units (MCUs), are widely used in the video conferencing field. Meanwhile, mainstream open-source streaming media servers based on the WebRTC protocol include Selective Forwarding Unit (SFU) servers, such as JANUS and MediaSoup from abroad, and SRS from China. All of these can solve the problem of media aggregation and distribution in the cloud.
[0102] Typically, as businesses grow, single-point services struggle to handle a large number of application access requests. To improve server capacity, cross-regional and cross-network capabilities, this naturally evolves into a streaming media service cluster that provides streaming media aggregation and forwarding in the cloud.
[0103] After forming a cluster using the SFU servers mentioned earlier, all incoming streaming media will be forwarded. While this offers good interactive features, large-scale distribution requires extensive forwarding within the streaming media cluster, resulting in high bandwidth consumption and cluster costs. Furthermore, the accumulated latency increases with each relay on the server. Therefore, employing appropriate technical measures to balance audio / video latency, server load, and server costs is necessary.
[0104] On the other hand, in traditional MCU solutions, audio and video decoding is performed on the streaming media server. From rendering / sampling to encoding, the server has high performance requirements and high costs. At the same time, it can only produce one or a few fixed media viewing layers, which is difficult to meet the personalized needs of the client. In addition, the mixing overhead of video on the MCU streaming media server will bring a lot of processing latency.
[0105] In other words, existing WebRTC and SFU solutions, which forward all participating media in full, put pressure on storage and network traffic. The capacity of a single server is limited, and forwarding scheduling is complex.
[0106] Limited by the display capabilities of the terminal, in the same display layer with multiple participants, in most scenarios it is not necessary to display the video of all participants, but the audio of all participants is required. Usually, audio and video are coupled to the same stream, so the client has to pull the video and audio of all participants. At this time, the downlink traffic of the terminal is large, and the server has to send a large amount of useless video data to the terminal, resulting in large traffic overhead.
[0107] II. Server Cluster
[0108] like Figure 2 The diagram shows the server cluster structure where the first edge server is located. The server cluster includes multiple edge servers and multiple central servers, as well as a message queue RabbitMQ server and a remote dictionary server (Redis) server. The cluster communicates internally via RabbitMQ and provides service interfaces to other business servers (such as MCU and Content Delivery Network (CDN) servers). The cluster status information is synchronized via Redis.
[0109] There is no physical difference between the central server and the edge server. The only distinction is based on their deployment location and network topology. Edge servers are typically deployed closer to the user and usually only have access to a single network, while central servers are deployed in the cloud and have cross-network capabilities that integrate multiple networks.
[0110] Both types of servers also have the capability to directly forward streaming media to publicly accessible internet services. This is used in specific scenarios to directly cascade with an MCU or CDN for distribution.
[0111] The audio mixing method, apparatus, edge server, and central server provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0112] like Figure 1 The diagram shown is a flowchart of one embodiment of the audio mixing method of this application. The method is applied to a first edge server, which is any edge server in a server cluster. The method includes:
[0113] Step 101: Obtain the client's business information;
[0114] It should be noted here that the client refers to the client that connects to each edge server in the server cluster to which the first edge server belongs; specifically, this step can be for the first edge server to obtain the load information of other edge servers by accessing Redis in the server cluster, such as clients that connect to other edge servers, in order to obtain the business information of clients that connect to each edge server.
[0115] In addition, the business information in this step may include the identifier of the edge server that the client is connected to, the client type, and the prior information that the user inputs on the client when the client connects to the edge server, such as the business type, business location range, number of participants, etc.
[0116] Step 102: Determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge mixing center-converging strategy, and full forwarding strategy;
[0117] This step determines the audio mixing strategy based on business information, enabling the execution of different audio mixing strategies for different business information. Specifically, employing an edge-converged mixing strategy to process the audio stream ensures a smooth access experience for traditional Session Initialization Protocol (SIP) conferencing devices.
[0118] The audio mixing method of this application embodiment obtains the client's business information and determines the audio mixing strategy corresponding to the business information. In this way, different audio mixing strategies are used to mix audio for different business information. Subsequently, the mixed audio is forwarded for interaction between clients. This solves the problems of high traffic overhead and high latency that exist when using full audio forwarding for interaction in the prior art.
[0119] As an optional implementation, step 102, determining the audio mixing strategy corresponding to the business information, includes:
[0120] Based on the business scenario determined by the aforementioned business information, the audio mixing strategy is determined; and / or
[0121] When a preset proportion of business participants are connected to the same edge server based on the business information, the audio mixing strategy is determined to be an edge convergence mixing strategy.
[0122] In other words, in this optional implementation, the client's business scenario can be determined based on business information, and the audio mixing strategy can be determined based on the business scenario. Alternatively, the audio mixing strategy corresponding to the business information can be determined based on the proportion of clients accessing each edge server.
[0123] As a specific implementation, determining the audio mixing strategy based on the business scenario identified by the business information includes any one of the following:
[0124] (1) When the business scenario is a scenario where the locations of business participants are concentrated, the audio mixing strategy is determined to be the edge convergence mixing strategy;
[0125] When business participants are located in a concentrated area, most of their clients may connect to the same edge server. In this case, an edge aggregation and mixing strategy can be used to mix the audio from each client. This approach saves on server cluster costs and meets the requirements for low-latency interaction.
[0126] (2) When the business scenario is a scenario where the number of business participants is greater than or equal to the first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy; wherein, by adopting this strategy, the load capacity of a single server can be reduced and high-quality, low-latency audio services can be provided.
[0127] On the one hand, the first preset value is a relatively large value, such as 50; on the other hand, the server cluster has the ability to access the nearest server. Therefore, if the business participants are located in different locations, different business participants will access different edge servers. In this case, if the edge aggregation and mixing strategy is adopted, the edge server with the edge mixing function enabled will be overloaded, which is not conducive to the transmission of audio stream. Therefore, in this case, the audio mixing strategy is determined to be a multi-edge mixing center aggregation strategy.
[0128] (3) When the business scenario is a scenario where the number of participants is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
[0129] The second preset value is a smaller value, such as 3. This means that in small-scale interactive scenarios, such as intercom and real-time communication scenarios, the full forwarding strategy is directly adopted.
[0130] This optional implementation selects different audio mixing strategies based on the number and concentration of business participants. When the load capacity of a single edge server is within its range, an edge aggregation mixing strategy is adopted to reduce the latency of audio sharing between different users. When the load capacity of a single edge server is exceeded, a multi-edge mixing center aggregation strategy is adopted to ensure that different users can share audio. When there are few business participants, the mixing function can be disabled and each audio stream can be forwarded directly, thus reducing server overhead.
[0131] Furthermore, as an optional implementation, the method also includes:
[0132] (1) When the audio mixing strategy is an edge convergence mixing strategy, if the edge mixing function of the first edge server is enabled, the first input audio stream and the second input audio stream received from at least one second edge server are mixed to obtain a full mixed stream; wherein, the first input audio stream is the audio stream of the client connected to the first edge server; and the second edge server is a server whose edge mixing function is not enabled.
[0133] It should be noted that the edge server with the edge mixing function enabled can be determined according to preset rules. For example, the preset rules can enable the edge mixing function for the edge server that has the first client access, and the preset rules can also enable the edge mixing function for the edge server with the most access clients.
[0134] When the audio mixing strategy is edge convergence mixing strategy and the first edge server has the edge mixing function enabled, the first edge server will request audio input stream signaling from other edge servers (second edge servers) through the RabbitMQ server in the server cluster. After obtaining the request audio input stream signaling by accessing the RabbitMQ server, each second edge server will send the audio input stream or (second input audio stream) of each client accessing the second edge server to the first edge server based on the User Data Protocol (UDP). After receiving each second input audio stream, the first edge server will mix the first input audio stream of each client accessing the first edge server with each second input audio stream to obtain a full mixed stream.
[0135] It should also be noted that, in this embodiment of the application, the client accessing the second edge server can send an audio stream or an audio / video stream to the second edge server. Therefore, the second edge server can send an audio stream or an audio / video stream to the first edge server. In the case of sending an audio / video stream, the first edge server can separate the audio / video stream to extract the audio stream.
[0136] Furthermore, this step can be specifically executed by the mixer of the mixing module of the first edge server. The specific mixing process can include: first, transcoding the received audio stream (lossy audio coding format (OPUS) audio data) into 16-bit Pulse Code Modulation (PCM) raw data; second, oversampling; and finally, superimposing and mixing. To prevent audio data overflow, this step uses a conventional adaptive thresholding algorithm for filtering.
[0137] (2) Filter the full mixed stream to obtain the output stream corresponding to each client connected to the first edge server;
[0138] It should be noted here that, considering the interactive capabilities of various client connections, such as when a client needs to specifically disable or adjust the voice of a particular participant, after mixing the received input audio streams, this step specifically filters the full mixed stream based on the needs of each connected client to obtain the output stream corresponding to each client. This step can be specifically executed by the filter in the mixing module of the first edge server.
[0139] (3) Send the full mixed stream to each of the second edge servers and send the output stream to the corresponding client.
[0140] This step involves sending the full mixed stream to each second edge server, enabling the second edge server to filter the full mixed stream based on the needs of each connected client, in order to generate an output stream that meets the needs of each client connected to the second edge server.
[0141] It's important to note that when using an edge-convergence hybrid strategy to process audio streams, on one hand, the server cluster needs to have proximity access capabilities. Proximity access is a common method in the server field, implemented using various technical means. In this application, for WebRTC protocol access, it mainly relies on the STUN protocol for IP address awareness to achieve proximity allocation. On the other hand, the server cluster is limited by the following: when the server's carrying capacity is insufficient and / or the number of clients accessing edge servers without edge-convergence functionality is too large, a smooth transition to a multi-edge-convergence center strategy is necessary to ensure the cluster's carrying capacity.
[0142] Furthermore, as an optional implementation, the method also includes:
[0143] Obtain the carrying capacity of each of the second edge servers;
[0144] In this step, the first edge server can query the carrying capacity of each second edge server by accessing the Redis server in the server cluster. The carrying capacity of the second edge server includes the number of clients accessing the second edge server.
[0145] Based on the CPU load of the first edge server and its carrying capacity, the audio mixing strategy is controlled to switch from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy.
[0146] In this optional implementation, the audio mixing strategy is switched from an edge-converging mixing strategy to a multi-edge mixing center-converging strategy based on the CPU load of the first edge server and the carrying capacity of the second edge server. This avoids the situation of insufficient server carrying capacity and achieves a balance between server carrying capacity and low latency.
[0147] As a specific implementation, controlling the audio mixing strategy to switch from the edge-converging mixing strategy to the multi-edge mixing center-converging strategy based on the CPU load of the first edge server and the carrying capacity includes:
[0148] The step of controlling the audio mixing strategy to switch from the edge-converging mixing strategy to the multi-edge mixing center-converging strategy based on the CPU load of the first edge server and the carrying capacity includes:
[0149] When it is determined that the CPU load has reached a first preset load value, and / or when it is sensed that the number of clients connected to at least one of the second edge servers has reached a first preset number of connections, the audio mixing strategy is controlled to smoothly transition from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy.
[0150] One specific implementation of this method is as follows: when the first edge server's carrying capacity is insufficient (CPU load reaches the first preset load value), more and more clients will join other edge servers (second edge servers). When the number of clients accessing another non-converging edge node approaches 50% of the server's carrying capacity, the audio mixing strategy is controlled to smoothly transition from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy.
[0151] This optional implementation ensures the carrying capacity of the server cluster by smoothly transitioning the audio mixing strategy from the edge-converging mixing strategy to the multi-edge-converging center-converging strategy.
[0152] As a more specific implementation, controlling the audio mixing strategy to smoothly transition from the edge-converging mixing strategy to the multi-edge-converging center-converging strategy includes:
[0153] The message queue server RabbitMQ is used to request the central server to start and notify the second edge server to start the edge hybrid function.
[0154] It should be noted that, after each server in the server cluster detects through Redis that the central server is turned on and the edge mixing function of each second edge server is turned on, it knows that it has smoothly transitioned to the multi-edge mixing center aggregation strategy. Each edge server (first edge server and second edge server) will mix the audio streams of the clients accessing it and send the mixed stream to the central server.
[0155] As an optional implementation, the method further includes:
[0156] When the audio mixing strategy is an edge convergence mixing strategy, if the first edge server does not enable the edge mixing function, the first input audio stream is sent to the third edge server so that the third edge server can perform audio mixing.
[0157] Wherein, the first input audio stream is the audio stream of the client connected to the first edge server, and the third edge server is a server with edge mixing function enabled.
[0158] In other words, when the current strategy is determined to be edge convergence mixing, if the first edge server detects that it has not enabled the edge mixing function, the first edge server can send the first input audio stream to the third edge server that has enabled the edge mixing function, so that the third edge server can perform audio mixing. The audio mixing process of the third edge server is similar to the process of mixing the audio stream when the first edge server enables the edge mixing function, and will not be described again here.
[0159] Furthermore, as an optional implementation, after the step of sending the first input audio stream to the third edge server, the method further includes:
[0160] Receive the full mixed stream sent by the third edge server;
[0161] The full mixed stream is filtered to obtain the output stream corresponding to each client connected to the first edge server;
[0162] The output stream is sent to the client corresponding to the output stream.
[0163] In other words, after receiving the full mixed stream from the third edge server, the first edge server can filter the full mixed stream based on the needs of each connected client to obtain an output stream that meets the needs of each client. In this way, the process of using the edge convergence mixing strategy to mix audio and output audio streams to each client is realized.
[0164] Optionally, when the audio mixing strategy is an edge-converging mixing strategy and the first edge server has not enabled edge mixing functionality, the process of smoothly transitioning from the edge-converging mixing strategy to a multi-edge mixing center-converging strategy includes:
[0165] Receive the signaling sent by the third server via RabbitMQ to enable the central server and enable the edge hybrid function;
[0166] Respond to this signaling to enable edge blending functionality;
[0167] The first input audio stream (the audio stream from the client connected to the first edge server) is mixed, and the mixed stream is sent to the central server.
[0168] Upon receiving the full mixed stream from the central server, stop sending the first input audio stream to the third edge server;
[0169] The full mixed stream sent by the central server is filtered to obtain the output stream corresponding to the client connected to the first edge server.
[0170] Optionally, when the audio mixing strategy is a multi-edge mixing center convergence strategy and the first edge server has not enabled edge mixing functionality, the process of smoothly transitioning from the multi-edge mixing center convergence strategy to the edge convergence mixing strategy includes:
[0171] The central server receives a fourth request message, which instructs the first edge server to disable the edge blending function and the sixth edge server to keep the edge blending function enabled.
[0172] In response to the fourth request message, disable the edge mixing function; receive the signaling from the sixth edge server requesting the first input audio stream;
[0173] In response to the signaling, the first input audio stream is sent to the sixth edge server so that the sixth edge server can mix the input audio streams received from each edge server and the audio streams from the clients accessing the sixth edge server;
[0174] Receive the full mixed stream sent by the sixth edge server;
[0175] The full mixed stream sent by the sixth edge server is filtered to obtain the output stream corresponding to each client connected to the first edge server.
[0176] Furthermore, as an optional implementation, the method also includes:
[0177] When the audio mixing strategy is a multi-edge mixing center convergence strategy, the first input audio streams of each client connected to the first edge server are mixed to obtain a first mixed stream;
[0178] Send a mixing request to the central server, the mixing request carrying the first mixed stream;
[0179] The central server receives a full mixed stream, wherein the full mixed stream is generated by the central server by mixing the received first mixed stream with a second mixed stream sent by other edge servers.
[0180] The full mixed stream is filtered to obtain the output stream corresponding to the client connected to the first edge server;
[0181] The output stream is sent to the client corresponding to the output stream.
[0182] The implementation process of this optional method enables the mixing of audio streams using a multi-edge convergence center mixing strategy, thus achieving dynamic interaction between clients while ensuring the server's capacity meets the requirements.
[0183] Furthermore, as an optional implementation, the method also includes:
[0184] (1) Receive a first request message sent by the central server, the request message being used to instruct the first edge server to keep the edge blending function enabled and to disable the edge blending function of other edge servers;
[0185] (2) In response to the first request message, keep the edge blending function in the enabled state so as to smoothly transition from the multi-edge blending center convergence strategy to the edge convergence blending strategy;
[0186] Wherein, the first request message is a message sent by the central server when it senses that the number of clients in the first channel has dropped to a third preset access number; the first channel is the channel where the clients accessing the first edge server are located; and the number of clients accessing the first edge server is the largest.
[0187] It should be noted that the first channel consists of various clients participating in the same service, their accessed edge servers, and the central server that initiates the service. In other words, a unique channel in the server cluster corresponds to a single multi-user interactive audio and video service for a particular service. Clients that join the same channel can access any audio data within the channel.
[0188] It should also be noted that when the number of clients in the first channel drops to the third preset access number, the central server senses that the load on the first channel has dropped to the number that a single edge can handle. Therefore, it is necessary to smoothly transition the audio mixing strategy to the edge convergence mixing strategy to reduce latency and server overhead.
[0189] Based on the above optional implementations, as a specific implementation, a smooth transition from the multi-edge hybrid center convergence strategy to the edge convergence hybrid strategy includes:
[0190] Receive a fourth input audio stream sent by other edge servers within the first channel;
[0191] The received fourth input audio stream and the first input audio stream are mixed;
[0192] Send the mixed stream to the other edge servers and stop sending the audio stream to the central server.
[0193] In other words, if it is determined that the audio mixing strategy has smoothly transitioned to the edge convergence mixing strategy, then the audio stream is stopped from being sent to the central server, thereby shutting down the central server and reducing latency and server overhead.
[0194] As another optional implementation, the method further includes:
[0195] The server receives a second request from the central server, which instructs the fourth edge server to keep the edge blending function enabled and disable the edge blending function of other edge servers.
[0196] In response to the second request, the edge blending function is disabled;
[0197] Receive a third request message from the fourth edge server requesting the first input audio stream;
[0198] In response to the third request message, the first input audio stream is sent to the fourth edge server;
[0199] Receive the mixed stream sent by the fourth edge server and stop sending the first input audio stream to the central server;
[0200] The fourth edge server is the server with the largest number of clients accessing the first channel, and the first channel is the channel where the clients accessing the first edge server are located.
[0201] It should be noted that this optional implementation is a process of smoothly transitioning from a multi-edge hybrid central aggregation strategy to an edge aggregation hybrid strategy when the number of clients connected to the first edge server is not the largest.
[0202] In this embodiment, the edge aggregation hybrid strategy and the multi-edge hybrid center aggregation strategy can be switched based on the load capacity (bearing capacity) of each server in the server cluster and the number of connected clients. This achieves a trade-off between server carrying capacity and low latency requirements. In particular, interactive audio and video services on windows-limited terminals will obtain a better user experience.
[0203] Here, the process of smoothly transitioning from a multi-edge hybrid center-convergence strategy to an edge-convergence hybrid strategy is explained:
[0204] (1) Logically promote the participating edge server to the central server, and other unselected edge servers forward the client's audio stream to the edge server.
[0205] (2) This edge server mixing module performs full mixing and pushes the mixed stream of this channel to other edge servers.
[0206] (3) After the edge server successfully receives and generates new real-time mixed audio, it stops sending audio streams to the central server.
[0207] (4) Reclaim the resources of the central server.
[0208] As a specific implementation, the full mixed stream is filtered to obtain the output stream corresponding to each client connected to the first edge server, including:
[0209] Based on the demand information of each client connected to the first edge server, a mute queue corresponding to each client is generated, and the mute queue is used to indicate audio streams that the client does not need;
[0210] In this step, the demand information for each client can be information pre-selected by the user when starting the client. For example, if the client is client B, the demand information can be that the audio stream from client A is not needed; then the mute queue includes clients A and B.
[0211] Based on the mute queue, the audio streams related to the mute queue in the full mix stream are filtered out to obtain the output stream.
[0212] It should be noted that the steps of this optional implementation can be specifically implemented by the filters in the mixing module of the first edge server. Furthermore, the various processes employing this optional implementation can meet the personalized needs of different clients.
[0213] As an optional implementation, the method further includes:
[0214] When the audio mixing strategy is a full forwarding strategy, a third input audio stream from a preset client is requested from other edge servers;
[0215] The received third input audio is sent to each client connected to the first edge server.
[0216] Below, in conjunction with Figure 3 The implementation process of the embodiments of this application will be described in detail below:
[0217] Before going into detail, let's first describe the structure of the edge server, which includes:
[0218] RabbitMQ module:
[0219] It is an industry-standard message sending and receiving module used for signaling interaction. Specifically, it accesses the RabbitMQ server through the RabbitMQ module to interact with other servers via signaling.
[0220] Streaming media forwarding control module:
[0221] The system uses UDP to receive and forward streaming media between the server and other servers or access clients. It separates the audio and video streams received by the client from the server and forwards them to other edge servers or mixing modules; in other words, the client provides an audio and video stream, and the edge server the client accesses separates the audio and video streams into audio and video streams.
[0222] Global state maintenance module:
[0223] The global status maintenance and statistics module provides dynamic status maintenance capabilities. First, it collects information about the current edge servers. This is achieved by periodically querying Redis for the load of other servers in the server cluster, as well as the status information of the streams and channels accessed by each server. The module then writes its own status and accessed stream and channel information to Redis for other services to query. Simultaneously, it synchronizes with other servers in the cluster via RabbitMQ when critical events occur, such as server online / offline status. Based on this, the cluster is provided with fault recovery and dynamic scaling capabilities.
[0224] Mixing module:
[0225] Considering compatibility with interactive capabilities across various access points, such as the need for clients to specifically disable or adjust the voice of a particular participant, and supporting cascading hybrid organization and scheduling, the mixing module differs somewhat from existing implementations.
[0226] The mixing module consists of two parts: a mixer and a filter.
[0227] The mixer is implemented by transcoding OPUS audio data into 16-bit PCM raw data, resampling, and then superimposing and mixing. To prevent audio data overflow, a conventional adaptive thresholding algorithm is used for filtering.
[0228] The implementation of the filter differs from conventional multi-party communication mixing in terms of output stream generation. In this application, the edge server needs to generate the current server's mixed stream. Then, the filter generates the receive stream required by a specific client. To accommodate the differentiated audio needs of each client, this application adds a silence queue to the filter. For example, when client A does not need client B's audio stream, B is added to A's silence queue. During filter runtime, the current audio data of B is first obtained through the silence queue and then filtered in the mixed stream output to A.
[0229] The execution process of this application embodiment is as follows:
[0230] Step 301: The client connects to the edge server;
[0231] Step 302: Determine whether audio mixing needs to be enabled based on the client's business information. If it is required, proceed to step 303 or step 307. If it is not required, proceed to step 315.
[0232] Step 303: Determine if the business scenario involves dispersed participants and a large number of participants, and implement a multi-edge hybrid central aggregation strategy;
[0233] Step 304: Enable the mixing module to mix the audio streams from the various clients that have connected;
[0234] Step 305: Forward the local mixed stream to the central server and request the full mixed stream from the central server;
[0235] Step 306: Filter the full mixed stream to remove the current client's own audio stream from the full mixed stream;
[0236] Step 307: Determine if business participants are concentrated in a specific region and implement an edge convergence hybrid strategy;
[0237] Step 308: Determine whether the current edge server has the hybrid module enabled; if yes, proceed to step 309; otherwise, proceed to step 312.
[0238] Step 309: Request other edge servers to push the audio stream of the specified client to this edge server via RabbitMQ;
[0239] Step 310: Create the mixing module for this channel;
[0240] Step 311: Send the output of the mixing module to the current client to the client, and send the full mix stream of the mixer to the audio source edge server;
[0241] Step 312: Forward the currently accessed audio stream to the aggregation edge server;
[0242] Step 313: Request the channel stream from the aggregation edge server, that is: request the full mixed stream from the aggregation edge server;
[0243] Step 314: The filter removes its own audio stream and sends the filtered audio stream to the current client;
[0244] Step 315: Determine if it is a low-participant business scenario (intercom, real-time communication);
[0245] Step 316: Request other edge servers to push the audio stream of the specified client to this edge server via RabbitMQ;
[0246] Step 317: Forward the audio stream to each client.
[0247] It should also be noted that since all servers in the server cluster (edge servers and central servers) communicate with each other via RabbitMQ, the audio stream is forwarded between the servers after they request the forwarding of the audio stream through signaling communication.
[0248] like Figure 4 As shown, this application embodiment provides an audio mixing method applied to a central server, including:
[0249] Step 401: Obtain the client's business information;
[0250] Similarly, this step can be achieved by the central server accessing the Redis server in the server cluster to obtain the business information of the clients connected to each edge server in the server cluster. The implementation process is similar to that of the first edge server, and will not be described in detail here.
[0251] Step 402: Determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge-converging center-converging strategy, and full-forwarding strategy.
[0252] The audio mixing method of this application embodiment obtains the client's business information and determines the audio mixing strategy corresponding to the business information. In this way, different audio mixing strategies are used to mix audio for different business information. Subsequently, the mixed audio is forwarded for interaction between clients. This solves the problems of high traffic overhead and high latency that exist when using full audio forwarding for interaction in the prior art.
[0253] As an optional implementation, step 402, determining the audio mixing strategy corresponding to the service information, includes:
[0254] Based on the business scenario determined by the aforementioned business information, the audio mixing strategy is determined; and / or
[0255] When a preset proportion of business participants are connected to the same edge server based on the business information, the audio mixing strategy is determined to be an edge convergence mixing strategy.
[0256] In other words, in this optional implementation, the client's business scenario can be determined based on business information, and the audio mixing strategy can be determined based on the business scenario. Alternatively, the audio mixing strategy corresponding to the business information can be determined based on the proportion of clients accessing each edge server.
[0257] As an optional implementation, determining the audio mixing strategy based on the business scenario identified by the business information includes any one of the following:
[0258] When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy;
[0259] When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy.
[0260] When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
[0261] This optional implementation selects different audio mixing strategies based on the number and concentration of business participants. When the load capacity of a single edge server is within the range, an edge aggregation mixing strategy is adopted to reduce the latency of audio sharing between different users; when the load capacity of a single edge server is exceeded, a multi-edge mixing center aggregation strategy is adopted to ensure that different users can share audio.
[0262] Furthermore, as an optional implementation, the method also includes:
[0263] When the audio mixing strategy is a multi-edge mixing center convergence strategy, the third mixed stream sent by each edge server is obtained;
[0264] The various third mixing streams are mixed to obtain the full mixing stream;
[0265] The full mixed stream is sent to each of the edge servers.
[0266] In other words, when the audio mixing strategy is a multi-edge mixing center aggregation strategy, each edge server mixes the input audio streams of the connected clients. When the RabbitMQ server receives a signal from the central server requesting the mixed streams from each edge server, it sends the mixed audio stream (third mixed stream) to the central server. The central server then mixes the received third mixed streams to obtain the full mixed stream. Finally, when the central server receives a signal from each edge server requesting the full mixed stream via the RabbitMQ service, it sends the full mixed stream to each edge server.
[0267] Furthermore, as an optional implementation, the method also includes:
[0268] When the audio mixing strategy is an edge-convergence mixing strategy, a request message to enable the central convergence function of the central server is received from the fifth edge server.
[0269] In response to the request message, the central aggregation function is enabled;
[0270] When the central aggregation function is enabled, it receives the mixed stream sent by each of the fifth edge servers;
[0271] The full mixed stream obtained by mixing the mixed streams sent by each of the fifth edge servers is sent to each of the edge servers.
[0272] In other words, if the audio mixing strategy is an edge-convergence mixing strategy, and a request message to enable the central convergence function of the central server is received, it is determined that the edge-convergence mixing strategy needs to be smoothly transitioned to a multi-edge-convergence central convergence strategy to ensure that audio sharing can be achieved between various clients.
[0273] It is important to emphasize that the forwarding of audio streams between the central server and edge servers, as well as the forwarding of audio streams between various edge servers, only occurs when the RabbitMQ server receives a signaling request for an audio stream from another server.
[0274] The beneficial effects of the audio mixing method in the embodiments of this application will be briefly explained below:
[0275] (1) The first edge server and the central server in the embodiment of this application are deployed in the same server cluster. The status of each node (server) within the cluster is synchronized through Redis to realize the separate processing of audio and video. It can be aggregated on any server within the cluster and has the function of pushing any channel selected streaming media (audio stream) to outside the cluster, as well as mixing audio streams from any channel. In addition, the edge server provides mixing capabilities to realize edge aggregation and mixing, meeting the needs of low-latency applications.
[0276] (2) By synchronizing Redis periodically in the global state maintenance module and synchronizing cluster information via real-time RabbitMQ messages, the state statistics of the media stream can be realized in real time, and specified data can be distributed to clients or other edge servers through the forwarding module. In its filter implementation, dynamic audio filtering is achieved by adding a mute list to each client to meet the needs of dynamic interaction.
[0277] (3) Server clusters adaptable to various scenarios in audio mixing organization and corresponding audio mixing methods that preserve interactivity. Based on the fundamental capabilities provided by the aforementioned two points, it responds to different scenario requirements. It is divided into the following three methods:
[0278] (a) Edge servers are not mixed; distribution is performed directly.
[0279] (b) Edge servers perform single edge aggregation hybridization;
[0280] (c) Central server aggregation and hybridization;
[0281] (b) and (c) are two types that can be dynamically adjusted smoothly based on business status. This achieves a trade-off between server capacity and low latency requirements. In particular, for window-limited terminal organizations providing interactive audio and video services, it saves a significant amount of downlink traffic, reduces cluster costs, and provides interactive performance, thus resulting in a better user experience.
[0282] (4) The audio processing in this application embodiment is based on channels. That is, each server that performs audio forwarding in this application embodiment is located in the same channel. This is beneficial for providing a standard reference audio stream when cascading to MCU servers and other third-party application servers. Audio and video synchronization is easier and easier to expand.
[0283] In summary, this application, based on the concept of audio and video separation (each edge server separates the audio and video streams from the clients accessing it), directly forwards video, while handling audio by adding mixing capabilities to the edge nodes of the streaming media service cluster. Mixing is enabled in scenarios requiring extremely low latency, and cascaded mixing is enabled when multi-edge convergence is needed. This approach can meet the needs of scenarios such as access from older devices, multi-user scenarios exceeding terminal display capabilities, interactive scenarios with storage or live streaming conversion, cloud-based chorus performances, and background music playback via live chat.
[0284] It should be noted that the audio mixing method provided in this application embodiment can be executed by an audio mixing device or a control module within the audio mixing device for executing the loading audio mixing method. This application embodiment uses the execution of the loading audio mixing method by an audio mixing device as an example to illustrate the audio mixing method provided in this application embodiment.
[0285] like Figure 5 As shown, this application embodiment provides an audio mixing device applied to a first edge server, comprising:
[0286] The first acquisition module 501 is used to acquire the client's business information;
[0287] The determining module 502 is used to determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the edge-converging mixing strategy, multi-edge mixing center-converging strategy, and full forwarding strategy.
[0288] The audio mixing device of this application embodiment has a first acquisition module 501 that acquires the client's service information and a determination module 502 that determines the audio mixing strategy corresponding to the service information. In this way, different audio mixing strategies are used for different service information to mix audio. Subsequently, the mixed audio is forwarded to facilitate interaction between clients. This solves the problems of high traffic overhead and high latency that exist when using full audio forwarding for interaction in the prior art.
[0289] Optionally, the determining module 502 includes:
[0290] The first determining submodule is used to determine the audio mixing strategy based on the business scenario determined by the business information; and / or
[0291] The second determining submodule is used to determine that the audio mixing strategy is an edge convergence mixing strategy when a preset proportion of business participants are connected to the same edge server based on the business information.
[0292] Optionally, the first determining submodule is specifically used to perform any of the following:
[0293] When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy;
[0294] When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy.
[0295] When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
[0296] Optionally, the device further includes:
[0297] The first mixing submodule is used to mix the first input audio stream and the second input audio stream received from at least one second edge server to obtain a full mixed stream when the audio mixing strategy is an edge convergence mixing strategy and the edge mixing function of the first edge server is enabled.
[0298] The filtering submodule is used to filter the full mixed stream to obtain the output stream corresponding to each client connected to the first edge server;
[0299] The first sending submodule is used to send the full mixed stream to each of the second edge servers and send the output stream to the corresponding client;
[0300] Wherein, the first input audio stream is the audio stream of the client connected to the first edge server, and the second edge server is a server whose edge mixing function is not enabled.
[0301] Optionally, the device further includes:
[0302] The second acquisition module is used to acquire the carrying capacity of each of the second edge servers;
[0303] The control module is used to control the audio mixing strategy to switch from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy based on the CPU load of the central processing unit of the first edge server and the carrying capacity.
[0304] Optionally, the control module is specifically configured to: when it is determined that the CPU load has reached a first preset load value, and / or when it is sensed that the number of clients connected to at least one of the second edge servers has reached a first preset number of connections, control the audio mixing strategy to smoothly transition from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy.
[0305] Optionally, when controlling the audio mixing strategy to smoothly transition from the edge-converging mixing strategy to the multi-edge mixing center-converging strategy, the control module is specifically used to: request the activation of the center server via the RabbitMQ message queue server and notify the second edge server to activate the edge mixing function.
[0306] Optionally, the device further includes:
[0307] The first sending module is configured to send a first input audio stream to the third edge server so that the third edge server can perform audio mixing if the first edge server has not enabled the edge mixing function when the audio mixing strategy is an edge convergence mixing strategy.
[0308] Wherein, the first input audio stream is the audio stream of the client connected to the first edge server, and the third edge server is a server with edge mixing function enabled.
[0309] Optionally, the device further includes:
[0310] The first receiving module is used to receive the full mixed stream sent by the third edge server;
[0311] The first filtering module is used to filter the full mixed stream to obtain the output stream corresponding to each client connected to the first edge server;
[0312] The second sending module is used to send the output stream to the client corresponding to the output stream.
[0313] Optionally, the device further includes:
[0314] The mixing module is used to mix the first input audio streams of each client connected to the first edge server to obtain a first mixed stream when the audio mixing strategy is a multi-edge mixing center convergence strategy;
[0315] The third sending module is used to send a hybrid request to the central server, the hybrid request carrying the first hybrid stream;
[0316] The second receiving module is used to receive the full mixed stream sent by the central server, wherein the full mixed stream is generated by the central server by mixing the received first mixed stream and the second mixed stream sent by other edge servers;
[0317] The second filtering module is used to filter the full mixed stream to obtain the output stream corresponding to the client connected to the first edge server;
[0318] The fourth sending module is used to send the output stream to the client corresponding to the output stream.
[0319] Optionally, the device further includes:
[0320] The third receiving module is used to receive a first request message sent by the central server. The first request message is used to instruct the first edge server to keep the edge blending function enabled and to disable the edge blending function of other edge servers.
[0321] The first response module is used to respond to the first request message by keeping the edge blending function in the enabled state so as to smoothly transition from the multi-edge blending center convergence strategy to the edge convergence blending strategy.
[0322] Wherein, the first request message is a message sent by the central server when it senses that the number of clients in the first channel has dropped to a third preset access number; the first channel is the channel where the clients accessing the first edge server are located; and the number of clients accessing the first edge server is the largest.
[0323] Optionally, the response module includes:
[0324] The receiving submodule is used to receive a fourth input audio stream sent by other edge servers within the first channel;
[0325] A mixing submodule is used to mix the received fourth input audio stream and the first input audio stream;
[0326] The second sending submodule is used to send the mixed stream to the other edge servers and stop sending the audio stream to the central server.
[0327] Optionally, the device further includes:
[0328] The fourth receiving module is used to receive a second request sent by the central server. The second request is used to instruct the fourth edge server to keep the edge mixing function enabled and to disable the edge mixing function of other edge servers.
[0329] The second response module is used to disable the edge blending function in response to the second request;
[0330] The fifth receiving module is used to receive a third request message sent by the fourth edge server for requesting the first input audio stream;
[0331] The third response module is used to send the first input audio stream to the fourth edge server in response to the third request message;
[0332] The processing module is used to receive the mixed stream sent by the fourth edge server and stop sending the first input audio stream to the central server;
[0333] The fourth edge server is the server with the largest number of clients accessing the first channel, and the first channel is the channel where the clients accessing the first edge server are located.
[0334] Optionally, when the first filtering module, the second filtering module, and the filtering submodule filter the full mixed stream to obtain the output stream corresponding to each client connected to the first edge server, they are specifically used for:
[0335] Based on the demand information of each client connected to the first edge server, a mute queue corresponding to each client is generated, and the mute queue is used to indicate audio streams that the client does not need;
[0336] Based on the mute queue, the audio streams related to the mute queue in the full mix stream are filtered out to obtain the output stream.
[0337] Optionally, the device further includes:
[0338] The fifth sending module is used to request a third input audio stream from a preset client from other edge servers when the audio mixing strategy is a full forwarding strategy;
[0339] The sixth receiving module is used to send the received third input audio to each client connected to the first edge server.
[0340] like Figure 6 As shown in the illustration, this application also provides an audio mixing device applied to a central server, comprising:
[0341] The first acquisition module 601 is used to acquire the client's business information;
[0342] The determining module 602 is used to determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the edge-converging mixing strategy, multi-edge mixing center-converging strategy, and full forwarding strategy.
[0343] The audio mixing device of this application embodiment has a first acquisition module 601 that acquires the client's service information and a determination module 602 that determines the audio mixing strategy corresponding to the service information. In this way, different audio mixing strategies are used for different service information to mix audio. Subsequently, the mixed audio is forwarded to facilitate interaction between clients. This solves the problems of high traffic overhead and high latency that exist when using full audio forwarding for interaction in the prior art.
[0344] Optionally, the determining module 602 includes:
[0345] The first determining submodule is used to determine the audio mixing strategy based on the business scenario determined by the business information; and / or
[0346] The second determining submodule is used to determine that the audio mixing strategy is an edge convergence mixing strategy when a preset proportion of business participants are connected to the same edge server based on the business information.
[0347] Optionally, the first determining submodule is specifically used to perform any of the following:
[0348] When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy;
[0349] When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy.
[0350] When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
[0351] Optionally, the device further includes:
[0352] The second acquisition module is used to acquire the third mixed stream sent by each edge server when the audio mixing strategy is a multi-edge mixing center convergence strategy;
[0353] A mixing module is used to mix the various third mixing streams to obtain the full mixing stream;
[0354] The first sending module is used to send the full mixed stream to each of the edge servers.
[0355] Optionally, the device further includes:
[0356] The first receiving module is configured to receive a request message from the fifth edge server to enable the central aggregation function of the central server when the audio mixing strategy is an edge aggregation mixing strategy.
[0357] The response module is used to respond to the request message and enable the central aggregation function;
[0358] The second receiving module is used to receive the mixed stream sent by each of the fifth edge servers when the central aggregation function is enabled.
[0359] The second sending module is used to send the full mixed stream obtained by mixing the mixed streams sent to each of the fifth edge servers to each of the edge servers.
[0360] Another embodiment of this application provides an edge server, such as... Figure 7 As shown, it includes a transceiver 710, a processor 700, a memory 720, and a program or instructions stored in the memory 720 and executable on the processor 700; when the processor 700 executes the program or instructions, it implements the audio mixing method applied to the first edge server as described above.
[0361] The transceiver 710 is used to receive and send data under the control of the processor 700.
[0362] Among them, Figure 7 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 700) and memory (memory 720). The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 710 can be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over a transmission medium. The processor 700 is responsible for managing the bus architecture and general processing, and the memory 720 can store data used by the processor 700 during operation.
[0363] This application embodiment also provides a central server, including: a transceiver, a processor, a memory, and a program or instructions stored in the memory and executable on the processor; when the processor executes the program or instructions, it implements the steps of the audio mixing method applied to the central server as described above.
[0364] The transceiver is used to receive and send data under the control of the processor.
[0365] The bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits represented by one or more processors (represented by processors) and memories (represented by memory). The bus architecture can also link various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides the interface. A transceiver can be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over a transmission medium. The processor is responsible for managing the bus architecture and general processing, and the memory can store data used by the processor during operation.
[0366] The structure of the central server is similar to that of the edge server.
[0367] This application also provides a readable storage medium storing a program. When executed by a processor, this program implements the various processes described above for the audio mixing method embodiment applied to a first edge server, or the various processes described above for the audio mixing method embodiment applied to a central server, achieving the same technical effect. To avoid repetition, these details are not repeated here. The readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0368] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An audio mixing method, characterized in that, Applied to the first edge server, including: Obtain the client's business information; Determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge mixing center-converging strategy, and full forwarding strategy; The step of determining the audio mixing strategy corresponding to the service information includes: Based on the business scenario determined by the aforementioned business information, the audio mixing strategy is determined; and / or, When a preset proportion of business participants are connected to the same edge server based on the business information, the audio mixing strategy is determined to be an edge convergence mixing strategy. The audio mixing strategy determined based on the business scenario identified by the business information includes any one of the following: When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy; When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy. When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
2. The method according to claim 1, characterized in that, The method further includes: When the audio mixing strategy is an edge convergence mixing strategy, if the edge mixing function of the first edge server is enabled, the first input audio stream and the second input audio stream received from at least one second edge server are mixed to obtain a full mixed stream. The full mixed stream is filtered to obtain the output stream corresponding to each client connected to the first edge server; The full mixed stream is sent to each of the second edge servers and the output stream is sent to the corresponding client. Wherein, the first input audio stream is the audio stream of the client connected to the first edge server, and the second edge server is a server whose edge mixing function is not enabled.
3. The method according to claim 2, characterized in that, The method further includes: Obtain the carrying capacity of each of the second edge servers; Based on the CPU load of the first edge server and its carrying capacity, the audio mixing strategy is controlled to switch from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy.
4. The method according to claim 3, characterized in that, The step of controlling the audio mixing strategy to switch from the edge-converging mixing strategy to the multi-edge mixing center-converging strategy based on the CPU load of the first edge server and the carrying capacity includes: When it is determined that the CPU load has reached a first preset load value, and / or when it is sensed that the number of clients connected to at least one of the second edge servers has reached a first preset number of connections, the audio mixing strategy is controlled to smoothly transition from the edge convergence mixing strategy to the multi-edge mixing center convergence strategy.
5. The method according to claim 4, characterized in that, Controlling the audio mixing strategy to smoothly transition from the edge-converging mixing strategy to the multi-edge-converging center-converging strategy includes: The message queue server RabbitMQ is used to request the central server to start and notify the second edge server to start the edge hybrid function.
6. The method according to claim 1, characterized in that, The method further includes: When the audio mixing strategy is an edge convergence mixing strategy, if the first edge server does not enable the edge mixing function, the first input audio stream is sent to the third edge server so that the third edge server can perform audio mixing. Wherein, the first input audio stream is the audio stream of the client connected to the first edge server, and the third edge server is a server with edge mixing function enabled.
7. The method according to claim 6, characterized in that, After the step of sending the first input audio stream to the third edge server, the method further includes: Receive the full mixed stream sent by the third edge server; The full mixed stream is filtered to obtain the output stream corresponding to each client connected to the first edge server; The output stream is sent to the client corresponding to the output stream.
8. The method according to claim 1, characterized in that, The method further includes: When the audio mixing strategy is a multi-edge mixing center convergence strategy, the first input audio streams of each client connected to the first edge server are mixed to obtain a first mixed stream; Send a mixing request to the central server, the mixing request carrying the first mixed stream; The central server receives a full mixed stream, wherein the full mixed stream is generated by the central server by mixing the received first mixed stream with a second mixed stream sent by other edge servers. The full mixed stream is filtered to obtain the output stream corresponding to the client connected to the first edge server; The output stream is sent to the client corresponding to the output stream.
9. The method according to claim 8, characterized in that, The method further includes: The first request message sent by the central server is received, which is used to instruct the first edge server to keep the edge blending function enabled and to disable the edge blending function of other edge servers. In response to the first request message, the edge blending function is kept on to smoothly transition from the multi-edge blending center convergence strategy to the edge convergence blending strategy; Wherein, the first request message is a message sent by the central server when it senses that the number of clients in the first channel has dropped to a third preset access number; the first channel is the channel where the clients accessing the first edge server are located; and the number of clients accessing the first edge server is the largest.
10. The method according to claim 9, characterized in that, The smooth transition from the multi-edge blending center convergence strategy to the edge convergence blending strategy includes: Receive a fourth input audio stream sent by other edge servers within the first channel; The received fourth input audio stream and the first input audio stream are mixed; Send the mixed stream to the other edge servers and stop sending the audio stream to the central server.
11. The method according to claim 8, characterized in that, The method further includes: The server receives a second request from the central server, which instructs the fourth edge server to keep the edge blending function enabled and disable the edge blending function of other edge servers. In response to the second request, the edge blending function is disabled; Receive a third request message from the fourth edge server requesting the first input audio stream; In response to the third request message, the first input audio stream is sent to the fourth edge server; Receive the mixed stream sent by the fourth edge server and stop sending the first input audio stream to the central server; The fourth edge server is the server with the largest number of clients accessing the first channel, and the first channel is the channel where the clients accessing the first edge server are located.
12. The method according to claim 2, 7 or 8, characterized in that, Filtering the full mixed stream to obtain the output stream corresponding to each client connected to the first edge server includes: Based on the demand information of each client connected to the first edge server, a mute queue corresponding to each client is generated, and the mute queue is used to indicate audio streams that the client does not need; Based on the mute queue, the audio streams related to the mute queue in the full mix stream are filtered out to obtain the output stream.
13. The method according to claim 1, characterized in that, The method further includes: When the audio mixing strategy is a full forwarding strategy, a third input audio stream from a preset client is requested from other edge servers; The received third input audio is sent to each client connected to the first edge server.
14. An audio mixing method, characterized in that, Applied to the central server, including: Obtain the client's business information; Determine the audio mixing strategy corresponding to the service information; wherein the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge mixing center-converging strategy, and full forwarding strategy; The step of determining the audio mixing strategy corresponding to the service information includes: Based on the business scenario determined by the aforementioned business information, the audio mixing strategy is determined; and / or When a preset proportion of business participants are connected to the same edge server based on the business information, the audio mixing strategy is determined to be an edge convergence mixing strategy. The audio mixing strategy determined based on the business scenario identified by the business information includes any one of the following: When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy; When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy. When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
15. The method according to claim 14, characterized in that, The method further includes: When the audio mixing strategy is a multi-edge mixing center convergence strategy, the third mixed stream sent by each edge server is obtained; The various third mixing streams are mixed to obtain a full mixing stream; The full mixed stream is sent to each of the edge servers.
16. The method according to claim 15, characterized in that, The method further includes: When the audio mixing strategy is an edge-convergence mixing strategy, a request message to enable the central convergence function of the central server is received from the fifth edge server. In response to the request message, the central aggregation function is enabled; When the central aggregation function is enabled, it receives the mixed stream sent by each of the fifth edge servers; The full mixed stream obtained by mixing the mixed streams sent by each of the fifth edge servers is sent to each of the edge servers.
17. An audio mixing device, characterized in that, Applied to the first edge server, including: The first acquisition module is used to acquire the client's business information; The determination module is used to determine the audio mixing strategy corresponding to the service information; wherein, the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge mixing center-converging strategy, and full forwarding strategy; The determining module includes: The first determining submodule is used to determine the audio mixing strategy based on the business scenario determined by the business information; and / or, The second determining submodule is used to determine that the audio mixing strategy is an edge convergence mixing strategy when a preset proportion of business participants are connected to the same edge server based on the business information. Specifically, the first determining submodule is used to perform any one of the following: When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy; When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy. When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
18. An audio mixing device, characterized in that, Applied to the central server, including: The first acquisition module is used to acquire the client's business information; The determination module is used to determine the audio mixing strategy corresponding to the service information; wherein, the audio mixing strategy includes any one of the following: edge-converging mixing strategy, multi-edge mixing center-converging strategy, and full forwarding strategy; The determining module includes: The first determining submodule is used to determine the audio mixing strategy based on the business scenario determined by the business information; and / or The second determining submodule is used to determine that the audio mixing strategy is an edge convergence mixing strategy when a preset proportion of business participants are connected to the same edge server based on the business information. Specifically, the first determining submodule is used to perform any one of the following: When the business scenario is one in which the locations of the business participants are concentrated, the audio mixing strategy is determined to be the edge-converging mixing strategy; When the business scenario is a scenario where the number of business participants is greater than or equal to a first preset value and the business participants are located in dispersed locations, the audio mixing strategy is determined to be the multi-edge mixing center convergence strategy. When the number of participants in the business scenario is less than the second preset value, the audio mixing strategy is determined to be the full forwarding strategy.
19. An edge server, comprising: A transceiver, a processor, a memory, and a program or instructions stored in the memory and executable on the processor; characterized in that, when the processor executes the program or instructions, it implements the steps of the audio mixing method as described in any one of claims 1 to 13.
20. A central server, comprising: A transceiver, a processor, a memory, and a program or instructions stored in the memory and executable on the processor; characterized in that, when the processor executes the program or instructions, it implements the steps of the audio mixing method as described in any one of claims 14 to 16.
21. A readable storage medium, characterized in that, The readable storage medium stores a program that, when executed by a processor, implements the steps of the audio mixing method as described in any one of claims 1 to 13, or the steps of the audio mixing method as described in any one of claims 14 to 16.
Citation Information
Patent Citations
Voice processing method, device and system
CN103327014A