Audio data processing method, storage medium and electronic device
By using an adaptive network mixing server to adaptively adjust the uplink and downlink network conditions and dynamically adjust the bit rate and sampling rate, the problem of smooth and high-quality audio data transmission in weak network environments is solved, and the effective processing of full-band audio data is achieved, improving transmission efficiency and quality.
Patent Information
- Application Number
- CN202211090684.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-09-07
AI Technical Summary
Existing technologies suffer from poor audio data transmission smoothness in weak network environments and cannot effectively process full-band audio data, resulting in low transmission efficiency and poor quality.
An adaptive network mixing server adaptively adjusts the uplink and downlink network status, dynamically adjusting the bit rate and sampling rate to achieve full-band audio data mixing.
It improves the smoothness and quality of audio data transmission in weak network environments, and enhances the transmission efficiency and quality of audio data under normal network conditions.
Smart Images

Figure CN116312579B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular, to an audio data processing method, a storage medium and an electronic device. BACKGROUND
[0002] At present, audio and video conference and audio and video communication are applied more and more widely in work and life scenes, and the transmission efficiency and transmission quality of audio data directly affect the work efficiency and user experience in related scenes, therefore, how to improve the transmission efficiency and transmission quality of audio data becomes one of the important problems in related fields.
[0003] In related technologies, the transmission efficiency and transmission quality of audio data are usually improved by a network mixing processing method. However, the existing scheme can only realize mixing processing of low-frequency band audio data, and cannot transmit and process full-frequency band audio data, and the existing technology has poor fluency in transmitting audio data in a weak network environment (such as low bandwidth, high delay, packet loss, etc.).
[0004] At present, no effective solution has been proposed for the above problems. SUMMARY
[0005] Embodiments of the present application provide an audio data processing method, a storage medium and an electronic device, to at least solve the technical problem that the audio data processing method provided by related technologies only processes low-frequency band audio in a normal network state, resulting in low transmission efficiency and poor transmission quality of audio data.
[0006] According to an aspect of an embodiment of the present application, an audio data processing method is provided, comprising: receiving multiple first audio data from multiple first terminals, wherein the audio processing parameters of each first audio data are adaptively adjusted based on the uplink network state between each first terminal and an adaptive network mixing server; mixing processing the multiple first audio data to obtain second audio data; and sending the second audio data to multiple second terminals, wherein the audio processing parameters of the second audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second terminal.
[0007] According to another aspect of the embodiments of the present application, there is also provided an audio data processing method, comprising: receiving a plurality of first online conference audio data from a plurality of first online conference terminals, wherein an audio processing parameter of each of the first online conference audio data is adaptively adjusted based on an uplink network state between each of the first online conference terminals and an adaptive network mixing server; mixing the first online conference audio data to obtain second online conference audio data; and sending the second online conference audio data to a plurality of second online conference terminals, wherein an audio processing parameter of the second online conference audio data is adaptively adjusted based on a downlink network state between the adaptive network mixing server and each of the second online conference terminals.
[0008] According to another aspect of the embodiments of the present application, there is also provided an audio data processing method, comprising: receiving a plurality of first online conference audio data from a plurality of first online conference terminals, wherein an audio processing parameter of each of the first online conference audio data is adaptively adjusted based on an uplink network state between each of the first online conference terminals and an adaptive network mixing server; mixing the first online conference audio data to obtain second online conference audio data; and sending the second online conference audio data to a plurality of second online conference terminals, wherein an audio processing parameter of the second online conference audio data is adaptively adjusted based on a downlink network state between the adaptive network mixing server and each of the second online conference terminals.
[0009] According to another aspect of the embodiments of the present application, there is also provided a computer readable storage medium comprising a stored program, wherein the program, when executed, controls a device in which the computer readable storage medium is located to perform any of the audio data processing methods described above.
[0010] According to another aspect of the embodiments of the present application, there is also provided an electronic device, comprising: a processor; and a memory connected to the processor and configured to provide the processor with instructions to perform the following processing steps: receiving a plurality of first audio data from a plurality of first terminals, wherein an audio processing parameter of each of the first audio data is adaptively adjusted based on an uplink network state between each of the first terminals and an adaptive network mixing server; mixing the first audio data to obtain second audio data; and sending the second audio data to a plurality of second terminals, wherein an audio processing parameter of the second audio data is adaptively adjusted based on a downlink network state between the adaptive network mixing server and each of the second terminals.
[0011] In the embodiment of the present application, multiple first audio data from multiple first terminals are received, wherein the audio processing parameters of each first audio data are adaptively adjusted based on the uplink network state between each first terminal and the adaptive network mixing server, the second audio data is obtained by mixing the multiple first audio data, and the second audio data is further sent to multiple second terminals, wherein the audio processing parameters of the second audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second terminal.
[0012] It is easily noticed that, by the embodiment of the present application, the uplink and downlink network states are adaptively adjusted based on the adaptive network mixing server to mix and transmit the audio data, achieving the purpose of receiving and sending multiple audio data by the adaptive network mixing server, thereby realizing the technical effect of improving the smoothness and quality of audio data transmission, and further solving the technical problem of the audio data processing method provided by the related art which only processes low frequency band audio in normal network state, resulting in low transmission efficiency and poor transmission quality. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings, which are included to provide a further understanding of the present application and constitute a part of this application, illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:
[0014] Figure 1 Fig. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing the audio data processing method;
[0015] Figure 2 Fig. 2 is a flowchart of an audio data processing method according to an embodiment of the present application;
[0016] Figure 3 Fig. 3 is a schematic diagram of an adaptive network logic according to an embodiment of the present application;
[0017] Figure 4 Fig. 4 is a schematic diagram of a mixing service logic according to an embodiment of the present application;
[0018] Figure 5 Fig. 5 is a flowchart of another audio data processing method according to an embodiment of the present application;
[0019] Figure 6 Fig. 6 is a flowchart of another audio data processing method according to an embodiment of the present application;
[0020] Figure 7 Fig. 7 is a structural schematic diagram of an audio data processing device according to an embodiment of the present application;
[0021] Figure 8 is a structural schematic diagram of an optional audio data processing device according to an embodiment of the present application;
[0022] Figure 9 is a structural schematic diagram of another audio data processing device according to an embodiment of the present application;
[0023] Figure 10 is a structural schematic diagram of another optional audio data processing device according to an embodiment of the present application;
[0024] Figure 11 is a structural schematic diagram of another audio data processing device according to an embodiment of the present application;
[0025] Figure 12 is a structural block diagram of another computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work should belong to the scope of protection of the present application.
[0027] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] First, some nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:
[0029] Adaptive Network Mixing Server (ANMS): refers to a server capable of mixing network video or network audio through an adaptive algorithm.
[0030] Web Real-Time Communications (WebRTC): can allow network applications or sites to establish a Peer-to-Peer link between browsers without the help of intermediaries, to realize the transmission of data stream (such as video stream, audio stream, etc.).
[0031] Embodiment 1
[0032] According to the embodiments of the present application, an audio data processing method embodiment is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0033] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the audio data processing method is shown. As shown in Figure 1 , the computer terminal 10 (or mobile device 10) can include one or more processors 102 (the processor 102 can include but not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports in the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can include more or less components than those shown in Figure 1 , or have a different configuration than that shown in Figure 1 .
[0034] It should be noted that the one or more processors 102 and / or other data processing circuits described above can be referred to herein as "data processing circuits" in general. The data processing circuit can be embodied in whole or in part as software, hardware, firmware or any other combination. In addition, the data processing circuit can be a single independent processing module, or all or part of any one of the other elements combined into the computer terminal 10 (or mobile device). As referred to in the embodiments of the present application, the data processing circuit serves as a processor control (for example, the selection of the variable resistance terminal path connected to the interface).
[0035] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the audio data processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implements the audio data processing method as described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include memories disposed remotely with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) which can be connected to other network devices through a base station so as to be able to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module which is used to communicate with the Internet in a wireless manner.
[0037] The display can be, for example, a touch screen type liquid crystal display (LCD) which can enable a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0038] It should be noted that, in some optional embodiments, the above-mentioned Figure 1 The computer device (or mobile device) shown can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that Figure 1 is merely one example of a particular implementation, and is intended to illustrate the types of components that can be present in the above-described computer device (or mobile device).
[0039] At present, audio and video conference and audio and video communication are applied more and more widely in work and life scenes. Taking a video conference as an example, the current speaker of the conference may be in a high packet loss, low bandwidth or high delay state of the uplink network, so that the voice of the current speaker appears to be stuck, unable to communicate, large delay, unable to understand the semantics of the speaker, and the like, which seriously affects the effect of the video conference. In addition, when the terminal of the current speaker sends audio data of a fixed code rate, the uplink network audio quality may be poor, and when the terminal of the conference listener receives audio data of a fixed code rate, the downlink network audio quality may also be poor. In particular, when the mixing efficiency of the mixing server is low, the full-band mixing may be stuck, which also affects the effect of the video conference.
[0040] In the related art, the transmission efficiency and quality of audio data are usually improved by a network mixing processing method (such as using a Freeswitch mixing service). Taking the use of the Freeswitch mixing service as an example, although the Freeswitch can realize the function of adaptive network mixing, the adaptive network mixing of the Freeswitch only realizes low-band mixing, cannot transmit and process full-band audio data, and the Freeswitch has poor smoothness in transmitting audio data in a weak network environment (such as low bandwidth, high delay, packet loss, and the like).
[0041] At present, no effective solution has been proposed for the above problems.
[0042] In the above running environment, the present application provides an audio data processing method as shown in Figure 2 . Figure 2 The flowchart of the audio data processing method according to an embodiment of the present application is shown in Figure 2 . The audio data processing method comprises the following steps.
[0043] Step S21, receiving multiple first audio data from multiple first terminals, wherein the audio processing parameters of each first audio data are adaptively adjusted based on the uplink network state between each first terminal and the adaptive network mixing server;
[0044] Step S22, mixing and processing the multiple first audio data to obtain second audio data;
[0045] Step S23, sending the second audio data to multiple second terminals, wherein the audio processing parameters of the second audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second terminal.
[0046] In the embodiments of the present application, the plurality of first terminals can be terminals that send audio data. The plurality of first audio data can be audio data sent by the plurality of first terminals. For example, in a video conference scenario, the plurality of first terminals can be terminal devices (such as a telephone, a smart phone, a microcontroller unit (MCU), a smart box, etc.) used by a plurality of conference speakers, and the plurality of first audio data can be speech audio data corresponding to the plurality of conference speakers.
[0047] Each of the plurality of first audio data corresponds to an audio processing parameter (such as a sampling rate, a code rate, etc.). The audio processing parameter can be adaptively adjusted based on an uplink network state between each first terminal and an adaptive network mix server (ANMS).
[0048] The plurality of first audio data is mixed to obtain second audio data. The mixing can integrate the plurality of first audio data into one audio track. Specifically, the frequency, dynamics, tone, positioning, reverberation, and sound field of each of the plurality of first audio data are adjusted separately, and then superimposed into one audio track to obtain the second audio data.
[0049] The plurality of second terminals can be terminals that receive audio data. For example, in a video conference scenario, the plurality of second terminals can be terminal devices (such as a telephone, a smart phone, an MCU, a smart box, etc.) used by a plurality of conference listeners.
[0050] During the process of sending the second audio data to the plurality of second terminals, the audio processing parameter (such as a sampling rate, a code rate, etc.) of the second audio data can be adaptively adjusted based on a downlink network state between the ANMS and each of the plurality of second terminals.
[0051] It should be noted that the plurality of first audio data can be full-band audio data, that is, the method provided by the embodiments of the present application can mix full-band audio data to obtain second audio data and send the second audio data to a plurality of second terminals.
[0052] Compared with the related art, the technical solution provided by the embodiments of the present application has the following technical innovations: the uplink and downlink network states (such as bandwidth evaluation data) are fed back to the ANMS for adaptive adjustment, such as dynamic adjustment of different code rates and sampling rates according to different network states; and full-band audio data can be mixed.
[0053] In the embodiment of the present application, multiple first audio data from multiple first terminals are received, wherein the audio processing parameters of each of the first audio data are adaptively adjusted based on the uplink network state between each of the first terminals and the adaptive network mixing server, the second audio data are obtained by mixing the multiple first audio data, and the second audio data are further transmitted to multiple second terminals, wherein the audio processing parameters of the second audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each of the second terminals.
[0054] It is easy to note that, by the embodiment of the present application, the uplink and downlink network states are adaptively adjusted based on the adaptive network mixing server to mix and transmit the audio data, so as to achieve the purpose of receiving and transmitting multiple audio data by the adaptive network mixing server, thereby realizing the technical effect of improving the smoothness and quality of audio data transmission, and further solving the technical problem of low transmission efficiency and poor transmission quality of the audio data processing method provided by the related art.
[0055] It should be noted that the embodiment of the present application can be but not limited to applied to the audio data processing application scenario, the mixing processing application scenario, the audio conference real-time processing application scenario, and the audio communication real-time processing application scenario. For example, it can also be applied to various application scenarios involving audio data processing in the technical fields of cloud computing, finance, transportation, manufacturing, energy, medical treatment (such as remote diagnosis, online consultation, etc.), education (such as network courses, teaching and research audio conference, online parent-teacher meeting, etc.), e-commerce (such as video live broadcast, audio live broadcast, product explanation, etc.), social communication (such as group video call, group voice call, group live broadcast, etc.), and the like.
[0056] In an optional embodiment, the audio data processing further comprises the following method steps:
[0057] In step S241, the uplink network state between each of the first terminals and the adaptive network mixing server is evaluated respectively to obtain a first evaluation result, wherein the first evaluation result comprises uplink network bandwidth evaluation, uplink network packet loss evaluation, and uplink network delay evaluation.
[0058] In step S242, the first evaluation result is fed back to the multiple first terminals, so that each of the first terminals adaptively adjusts the audio processing parameters of the corresponding first audio data by using the first evaluation result.
[0059] In the optional embodiment, the evaluating the uplink network state comprises evaluating the uplink network bandwidth, evaluating the uplink network packet loss and evaluating the uplink network delay. The first evaluation result can comprise the uplink network bandwidth evaluation, the uplink network packet loss evaluation and the uplink network delay evaluation. The audio processing parameter of the first audio data can comprise a sampling rate, a code rate and the like.
[0060] For example, in the audio call diagnosis scene in the medical field, the method provided by the embodiment of the application is used for real-time processing of audio data, and the transmission quality of the audio data is improved in the case of ensuring smooth voice communication for different uplink network states of different users (such as doctors or patients).
[0061] Figure 3 is a schematic diagram of an optional adaptive network logic according to the embodiment of the application, as shown in Figure 3 During the audio call diagnosis process: the terminal (such as a smart phone) used by the current speaker (such as a doctor or a patient) performs real-time network evaluation with the media mixing server; then, the terminal used by the current speaker dynamically adjusts the sending code rate of the audio data according to the uplink network evaluation data (including the bandwidth condition, the packet loss condition and the delay condition).
[0062] In an optional embodiment, the audio data processing further comprises the following method steps:
[0063] Step S251, respectively evaluating the uplink network state between each first terminal in the plurality of first terminals and the adaptive network mixing server to obtain a first evaluation result, wherein the first evaluation result comprises an uplink network bandwidth evaluation, an uplink network packet loss evaluation and an uplink network delay evaluation;
[0064] Step S252, adaptively adjusting the first network jitter balancing parameter corresponding to each first terminal in the plurality of first terminals based on the first evaluation result.
[0065] In the optional embodiment, the evaluating the uplink network state comprises evaluating the uplink network bandwidth, evaluating the uplink network packet loss and evaluating the uplink network delay. The first evaluation result can comprise the uplink network bandwidth evaluation, the uplink network packet loss evaluation and the uplink network delay evaluation. The first network jitter balancing parameter corresponding to each first terminal can be used for balancing the network jitter of the uplink network. The network jitter refers to the case that the data transmission on the network is fast or slow.
[0066] Still as shown in Figure 3 The uplink network state between the media mixing server and the speaker terminal is evaluated to obtain uplink network evaluation data, which can comprise uplink network bandwidth data, uplink network packet loss data and uplink network delay data.
[0067] Still as Figure 3 shown, according to the above uplink network evaluation data (reflecting the current network packet receiving situation), the media mixing server can adaptively adjust the network parameters (such as network jitter balancing parameters) to improve the stability, smoothness and real-time performance of network transmission of audio data, and reduce network jitter.
[0068] In an optional embodiment, the audio data processing further includes the following method steps:
[0069] Step S261, respectively evaluate the uplink network state between each first terminal in the plurality of first terminals and the adaptive network mixing server to obtain a first evaluation result, wherein the first evaluation result includes: uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation;
[0070] Step S262, based on the first evaluation result, limit the packet loss compensation data corresponding to each first terminal in the plurality of first terminals.
[0071] In the above optional embodiment, evaluating the uplink network state includes: evaluating the uplink network bandwidth, evaluating the uplink network packet loss and evaluating the uplink network delay. The above first evaluation result can include: uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation. The packet loss compensation data corresponding to each first terminal can be the uplink network corresponding packet loss retransmission (NACK) logic.
[0072] Still as Figure 3 shown, the uplink network state between the media mixing server and the speaker terminal is evaluated to obtain uplink network evaluation data, which can include: uplink network bandwidth data, uplink network packet loss data and uplink network delay data.
[0073] Still as Figure 3 shown, in order to avoid occupying a large uplink network bandwidth and a large CPU resource, according to the above uplink network evaluation data, the terminal used by the current speaker can process the packet loss retransmission (NACK) logic of the uplink network (equivalent to limit the packet loss compensation number) to achieve low nack packet.
[0074] In an optional embodiment, the audio data processing further includes the following method steps:
[0075] Step S271, respectively receive the second evaluation result from each second terminal in the plurality of second terminals, wherein the second evaluation result is used to evaluate the downlink network state between each second terminal and the adaptive network mixing server, and the second evaluation result includes: downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation;
[0076] At step S272, the audio processing parameter of the second audio data corresponding to each second terminal is adaptively adjusted based on the second evaluation result.
[0077] In the optional embodiment described above, the second evaluation result can include downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation. The downlink network state is evaluated by evaluating the downlink network bandwidth, evaluating the downlink network packet loss and evaluating the downlink network delay. The audio processing parameter of the second audio data can include the sampling rate, the code rate, etc.
[0078] Still taking the real-time processing of audio data by the method provided by the embodiment of the application in the audio call consultation scenario in the medical field as an example, for different downlink network states of different users (such as doctors or patients), the transmission quality of the audio data is improved while ensuring smooth call sound.
[0079] Still as shown in Figure 3 During the audio call consultation, the terminal (such as a smart phone) used by the current listener (such as a patient or a doctor) performs real-time network evaluation with the media mixing server. Then, the terminal used by the current listener dynamically adjusts the receiving encoding code rate of the audio data according to the downlink network evaluation data (including the bandwidth condition, the packet loss condition and the delay condition).
[0080] In an optional embodiment, in the audio data processing method, the second evaluation result is also used for adaptive adjustment of the second network jitter balancing parameter by each second terminal in the plurality of second terminals, and flow limiting of the packet loss compensation data by each second terminal in the plurality of second terminals.
[0081] In the optional embodiment described above, the second evaluation result can include downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation. The second network jitter balancing parameter corresponding to each second terminal can be used to balance the network jitter of the downlink network. The network jitter refers to the condition that the data transmission on the network is fast or slow. The packet loss compensation data corresponding to each second terminal can be the packet loss retransmission (NACK) logic of the downlink network.
[0082] Still as shown in Figure 3 The downlink network evaluation data (including downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation) obtained by the terminal used by the current listener for evaluating the downlink network is received.
[0083] The terminal used by the current listener can also dynamically adjust network parameters (such as network jitter balancing parameters) according to the above downlink network evaluation data, to ensure the smoothness and real-time performance of audio data transmission in the downlink network. In addition, in order to avoid occupying a large downlink network bandwidth and a large CPU resource, the terminal used by the current listener can process the downlink network packet retransmission (NACK) logic (equivalent to limiting the number of packet compensation), to achieve a low nack packet.
[0084] In an optional embodiment, in step S22, the mixing processing is performed on the plurality of first audio data to obtain the second audio data, including the following method steps:
[0085] In step S221, the mixing processing is performed on the plurality of first audio data to obtain the third audio data.
[0086] In step S222, the post-processing corresponding to each first terminal is performed on the third audio data to obtain the second audio data, wherein the post-processing includes at least one of the following: reverberation processing, gain adjustment.
[0087] In the above optional embodiment, the third audio data can be audio data on a single audio track after the mixing processing is performed on the plurality of first audio data.
[0088] Since different users have different requirements for receiving audio data, the above third audio data can be subjected to post-processing corresponding to each first terminal, and then the second audio data is obtained, which is the to-be-received audio data of the receiving terminal corresponding to each first terminal (equivalent to the above plurality of second terminals).
[0089] The above post-processing can include reverberation processing and gain adjustment. The reverberation processing can be adjustment of frequency, dynamics, sound quality, positioning, reverberation and sound field of the third audio data. The gain adjustment can be adjustment of a current signal, a voltage signal or a power signal corresponding to the third audio data (the adjustment is specified in decibels).
[0090] Figure 4 is a schematic diagram of an optional mixing service logic according to an embodiment of the present application, as Figure 4 shown, using the method provided by the embodiment of the present application, in the process of real-time processing of audio data based on the self-developed ANMS, the input end can be a telephone, an application, an MCN device (such as a microcontroller unit in Figure 4 ) and a smart box, etc.
[0091] Still as Figure 4As shown, different networks, different encoders and different code rates can be adapted for different terminals corresponding to different input terminals, and the downlink code rate is modified in real time for network adaptation. Different pre-processing and post-processing can be performed for different users, thereby improving the audio interactive experience of users.
[0092] Specifically, still as Figure 4 As shown, after real-time transmission protocol packets, jitter buffer processing, decoder processing, network equalizer processing and pre-processing (optional) for different users, mixed audio data (equivalent to the third audio data described above) can be obtained through mixer mixing processing. Further, according to the receiving requirements of different users, different post-processing can be performed on the mixed audio data, and then through encoder processing and real-time transmission protocol packets, the audio data to be received can be obtained. The audio data to be received can be obtained by processing all the audio data input corresponding to the user.
[0093] It is easy to note that, compared with the related art, the technical innovation point of the technical solution provided by the present application is that the bandwidth evaluation data of the uplink and downlink network states is fed back to the audio encoder for adaptive adjustment, such as dynamic adjustment of different code rate encoding and sampling rate according to different network states; and the mixing processing of full-band audio data is realized based on WebRTC.
[0094] It should be noted that the prior art (using Freeswitch mixing service) usually uses 16K sampling rate under G711 encoding of the telephone to perform low-band mixing, compared with the prior art, the method provided by the embodiment of the present application can solve the full-band mixing and high and low code rate encoding under 48K sampling rate based on self-developed ANMS, thereby obtaining smoother and higher quality audio.
[0095] In an optional embodiment, a graphical user interface is provided by the adaptive network mixing server, and the content displayed by the graphical user interface at least partially contains an audio processing scene, and the audio data processing method further includes the following method steps:
[0096] Step S281, in response to a first touch operation acting on the graphical user interface, setting a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger the adaptive adjustment of the audio processing parameters of each first audio data based on the uplink network state for each first terminal, and the second parameter adjustment option is used to trigger the adaptive adjustment of the audio processing parameters of the second audio data based on the downlink network state for the adaptive network mixing server;
[0097] Step S282, in response to a second touch operation acting on the graphical user interface, selecting multiple first audio data from the candidate audio data set;
[0098] In response to the third touch operation on the graphical user interface, the multiple first audio data are mixed to obtain second audio data.
[0099] In response to the fourth touch operation on the graphical user interface, the second audio data are sent to the multiple second terminals.
[0100] In the optional embodiment, the adaptive network mixing server can be an ANMS developed by the user. A graphical user interface can be provided by the ANMS. The graphical user interface can display an audio processing scene.
[0101] In the optional embodiment, the user can perform a first touch operation on the graphical user interface. The user can touch a "setting" button or a "parameter adjustment option" button in the graphical user interface, or perform a specified setting touch gesture in a preset parameter adjustment touch area, to set a first parameter adjustment option and a second parameter adjustment option. The first parameter adjustment option is used to trigger each first terminal to adaptively adjust the audio processing parameter of each first audio data based on the uplink network state. The second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameter of the second audio data based on the downlink network state.
[0102] In the optional embodiment, the user can also perform a second touch operation on the graphical user interface. The user can touch a "select" button or a "confirm" button in the graphical user interface, or perform a point selection, a frame selection, or a long press on the candidate audio data set displayed in the graphical user interface, to select the multiple first audio data from the candidate audio data set.
[0103] In the optional embodiment, the user can also perform a third touch operation on the graphical user interface. The user can touch a "mixing" button, a "processing" button, or a "confirm" button in the graphical user interface, to mix the multiple first audio data to obtain the second audio data.
[0104] In the optional embodiment, the user can also perform a fourth touch operation on the graphical user interface. The user can touch a "send" button or a "confirm" button in the graphical user interface, to send the second audio data to the multiple second terminals.
[0105] Specifically, the first, second, third, and fourth touch operations described above can all be operations performed by a user touching the display screen of the terminal device with their finger and interacting with the terminal device. These touch operations can include single-point touch and multi-point touch, where the touch operation for each touch point can include clicking, long-pressing, hard-pressing, swiping, etc. The first, second, third, and fourth touch operations can also be touch operations implemented through input devices such as a mouse or keyboard.
[0106] In summary, the key features of this invention are: improving the smoothness of audio transmission in weak network environments by leveraging the characteristics of the ANMS model, such as resistance to packet loss, low bandwidth, and high latency; and enhancing the quality of audio (full-bandwidth) transmission in normal network and high-bandwidth environments through the mixing function based on the ANMS model.
[0107] Therefore, the technical problem solved by the above-mentioned method provided by the embodiments of the present invention is at least that: it can improve the smoothness of audio data transmission under the conditions of high packet loss, low bandwidth and high latency in the uplink network state, and can also improve the quality of audio data transmission under normal conditions (normal bandwidth or high bandwidth) in the uplink network state or downlink network state.
[0108] The beneficial effects of the technical solution provided by the present invention are at least as follows: it can improve the anti-packet loss capability, low bandwidth transmission capability and high latency transmission capability of audio data transmission in weak network environments, thereby improving the smoothness of audio data transmission in weak network environments.
[0109] Under the above operating environment, the present invention provides, as follows: Figure 5 This illustrates an audio data processing method. Figure 5 This is a flowchart of another audio data processing method according to an embodiment of the present invention, such as... Figure 5 As shown, the audio data processing method includes:
[0110] Step S51: Receive multiple first online conference audio data from multiple first online conference terminals, wherein the audio processing parameters of each first online conference audio data are adaptively adjusted based on the uplink network status between each first online conference terminal and the adaptive network mixing server.
[0111] Step S52: Mix the multiple first online conference audio data to obtain the second online conference audio data;
[0112] Step S53: Send the second online conference audio data to multiple second online conference terminals. The audio processing parameters of the second online conference audio data are adaptively adjusted based on the downlink network status between the adaptive network mixing server and each second online conference terminal.
[0113] In the embodiments of the present application, the plurality of first online conference terminals can be terminals that send audio data in the online conference process. The plurality of first online conference audio data can be online conference audio data sent by the plurality of first online conference terminals. The plurality of first online conference terminals can be terminal devices (such as telephones, smart phones, MCUs, smart boxes, etc.) used by a plurality of online conference speakers, and the plurality of first online conference audio data can be speech audio data corresponding to the plurality of online conference speakers.
[0114] Each of the plurality of first online conference audio data corresponds to an audio processing parameter (such as a sampling rate, a code rate, etc.). The audio processing parameter can be adaptively adjusted based on the uplink network state between each first online conference terminal and the adaptive network mixing server (ANMS).
[0115] The plurality of first online conference audio data is mixed to obtain second online conference audio data. The mixing can integrate the plurality of first online conference audio data into one audio track. Specifically, the frequency, dynamics, sound quality, positioning, reverberation, and sound field of each of the plurality of first online conference audio data are adjusted separately, and then superimposed into one audio track to obtain the second online conference audio data.
[0116] The plurality of second online conference terminals can be terminals that receive audio data in the online conference process. The plurality of second online conference terminals can be terminal devices (such as telephones, smart phones, MCUs, smart boxes, etc.) used by a plurality of online conference listeners.
[0117] During the process of sending the second online conference audio data to the plurality of second online conference terminals, the audio processing parameter (such as the sampling rate, the code rate, etc.) of the second online conference audio data can be adaptively adjusted based on the downlink network state between the ANMS and each of the plurality of second online conference terminals.
[0118] It should be noted that the plurality of first online conference audio data can be online conference full-band audio data, that is, the method provided by the embodiments of the present application can mix the online conference full-band audio data to obtain second online conference audio data and send it to a plurality of second online conference terminals.
[0119] Compared with the related art, the technical solution provided by the embodiments of the present application has the following technical innovations: the uplink and downlink network states (such as bandwidth evaluation data) in the online conference process are fed back to the ANMS for adaptive adjustment, such as dynamic adjustment of different code rates and sampling rates according to different network states; and the mixing of online conference full-band audio data can be realized.
[0120] In the embodiment of the present application, multiple first online conference audio data from multiple first online conference terminals are received, wherein the audio processing parameters of each of the first online conference audio data are adaptively adjusted based on the uplink network status between each of the first online conference terminals and the adaptive network mixing server, the multiple first online conference audio data are mixed to obtain second online conference audio data, and the second online conference audio data are further transmitted to multiple second online conference terminals, wherein the audio processing parameters of the second online conference audio data are adaptively adjusted based on the downlink network status between the adaptive network mixing server and each of the second online conference terminals.
[0121] It is easy to note that, through the embodiment of the present application, the uplink and downlink network status during the online conference is adaptively adjusted based on the adaptive network mixing server to mix and transmit the online conference audio data, so as to achieve the purpose of receiving and transmitting multiple online conference audio data through the adaptive network mixing server, thereby realizing the technical effect of improving the smoothness and quality of online conference audio data transmission, and further solving the technical problem that the audio data processing method provided by the related art only processes low-frequency band audio in normal network status, resulting in low transmission efficiency and poor transmission quality.
[0122] In an optional embodiment, a graphical user interface is provided through the adaptive network mixing server, and the content displayed by the graphical user interface at least partially contains an online conference audio processing scene, and the audio data processing method further includes the following method steps:
[0123] Step S541, in response to a first touch operation acting on the graphical user interface, setting a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each of the first online conference terminals to adaptively adjust the audio processing parameters of each of the first online conference audio data based on the uplink network status, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameters of the second online conference audio data based on the downlink network status;
[0124] Step S542, in response to a second touch operation acting on the graphical user interface, selecting multiple first online conference audio data from the candidate online conference audio data set;
[0125] Step S543, in response to a third touch operation acting on the graphical user interface, mixing the multiple first online conference audio data to obtain second online conference audio data;
[0126] In response to the fourth touch operation on the graphical user interface, the second online conference audio data is sent to the plurality of second online conference terminals.
[0127] In the optional embodiments described above, the adaptive network mixing server can be a self-developed ANMS. A graphical user interface can be provided through the ANMS. The graphical user interface can display a wired online conference audio processing scene.
[0128] In the optional embodiments described above, the user can perform a first touch operation on the graphical user interface. The user can touch a "setting" button or a "parameter adjustment option" button in the graphical user interface, or perform a specified setting touch gesture in a preset parameter adjustment touch area, to set a first parameter adjustment option and a second parameter adjustment option, where the first parameter adjustment option is used to trigger each first online conference terminal to adaptively adjust the audio processing parameter of each piece of first online conference audio data based on the uplink network state, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameter of the second online conference audio data based on the downlink network state.
[0129] In the optional embodiments described above, the user can also perform a second touch operation on the graphical user interface. The user can touch a "select" button or a "confirm" button in the graphical user interface, or perform a point selection, a box selection, or a long press on the candidate audio data set displayed in the graphical user interface, to select a plurality of pieces of first online conference audio data from the candidate online conference audio data set.
[0130] In the optional embodiments described above, the user can also perform a third touch operation on the graphical user interface. The user can touch a "mixing" button, a "processing" button, or a "confirm" button in the graphical user interface, to perform mixing processing on the plurality of pieces of first online conference audio data to obtain the second online conference audio data.
[0131] In the optional embodiments described above, the user can also perform a fourth touch operation on the graphical user interface. The user can touch a "send" button or a "confirm" button in the graphical user interface, to send the second online conference audio data to the plurality of second online conference terminals.
[0132] Specifically, the first, second, third, and fourth touch operations described above can all be operations performed by a user touching the display screen of the terminal device with their finger and interacting with the terminal device. These touch operations can include single-point touch and multi-point touch, where the touch operation for each touch point can include clicking, long-pressing, hard-pressing, swiping, etc. The first, second, third, and fourth touch operations can also be touch operations implemented through input devices such as a mouse or keyboard.
[0133] In summary, the key features of this invention are: based on the characteristics of the ANMS model, such as resistance to packet loss, low bandwidth, and high latency, to improve the smoothness of online conference audio transmission in weak network environments; and based on the mixing function of the ANMS model, to improve the quality of online conference audio (full bandwidth) transmission in normal network and high bandwidth environments.
[0134] Therefore, the technical problem solved by the above-mentioned method provided by the embodiments of the present invention is at least as follows: it can improve the smoothness of online conference audio data transmission under the conditions of high packet loss, low bandwidth and high latency in the uplink network state of online conferences, and can also improve the quality of online conference audio data transmission under normal conditions (normal bandwidth or high bandwidth) in the uplink or downlink network state of online conferences.
[0135] The beneficial effects of the technical solution provided by the present invention are at least as follows: it can improve the anti-packet loss capability, low bandwidth transmission capability and high latency transmission capability of online conference audio data transmission in weak network environments, thereby improving the smoothness of online conference audio data transmission in weak network environments.
[0136] Under the above operating environment, the present invention provides, as follows: Figure 6 This illustrates an audio data processing method. Figure 6 This is a flowchart of another audio data processing method according to an embodiment of the present invention, such as... Figure 6 As shown, the audio data processing method includes:
[0137] Step S61: Receive multiple channels of first online classroom audio data from multiple first online classroom terminals, wherein the audio processing parameters of each channel of first online classroom audio data are adaptively adjusted based on the uplink network status between each first online classroom terminal and the adaptive network mixing server.
[0138] Step S62: Mix the multiple first online classroom audio data to obtain second online classroom audio data;
[0139] Step S63, sending the second online class audio data to the plurality of second online class terminals, wherein the audio processing parameters of the second online class audio data are adaptively adjusted based on the downlink network status between the adaptive network mixing server and each second online class terminal.
[0140] In the embodiments of the present application, the plurality of first online class terminals can be terminals that send audio data in the online class process. The above-mentioned multi-channel first online class audio data can be online class audio data sent by the above-mentioned plurality of first online class terminals. The plurality of first online class terminals can be terminal devices (such as telephones, smart phones, MCUs, smart boxes, etc.) used by a plurality of online class speakers (such as teachers), and the multi-channel first online class audio data can be speaking audio data corresponding to the plurality of online class speakers.
[0141] Each of the above-mentioned multi-channel first online class audio data corresponds to audio processing parameters (such as sampling rate, code rate, etc.). The audio processing parameters can be adaptively adjusted based on the uplink network status between each first online class terminal and the adaptive network mixing server (ANMS).
[0142] The above-mentioned multi-channel first online class audio data is mixed to obtain second online class audio data. The mixing process can integrate the multi-channel first online class audio data into one audio track. Specifically, the frequency, dynamics, sound quality, positioning, reverberation and sound field of each channel of the multi-channel first online class audio data are adjusted separately, and then superimposed into one audio track, thereby obtaining the above-mentioned second online class audio data.
[0143] The above-mentioned plurality of second online class terminals can be terminals that receive audio data in the online class process. The plurality of second online class terminals can be terminal devices (such as telephones, smart phones, MCUs, smart boxes, etc.) used by a plurality of online class listeners (such as a plurality of students).
[0144] During the process of sending the above-mentioned second online class audio data to the above-mentioned plurality of second online class terminals, the audio processing parameters (such as sampling rate, code rate, etc.) of the second online class audio data can be adaptively adjusted based on the downlink network status between the ANMS and each second online class terminal of the plurality of second online class terminals.
[0145] It should be noted that the above-mentioned multi-channel first online class audio data can be online class full-band audio data, that is, the above-mentioned method provided by the embodiments of the present application can mix the online class full-band audio data to obtain second online class audio data and send it to a plurality of second online class terminals.
[0146] Compared with the related art, the technical solution provided by the embodiment of the application has the following technical innovation points: the uplink and downlink network states (such as bandwidth evaluation data) in the online classroom process are fed back to the ANMS for adaptive adjustment, such as dynamic adjustment of different code rate coding and sampling rate according to different network states; and the online classroom full-band audio data can be mixed and processed.
[0147] In the embodiment of the application, multi-path first online classroom audio data from a plurality of first online classroom terminals is received, wherein the audio processing parameters of each path of the first online classroom audio data are adaptively adjusted based on the uplink network state between each first online classroom terminal and the adaptive network mixing server, the second online classroom audio data is obtained by mixing and processing the multi-path first online classroom audio data, and the second online classroom audio data is further transmitted to a plurality of second online classroom terminals, wherein the audio processing parameters of the second online classroom audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second online classroom terminal.
[0148] It is easy to note that, by the embodiment of the application, the uplink and downlink network states in the online classroom process are adaptively adjusted based on the adaptive network mixing server to mix and process the online classroom audio data and transmit the online classroom audio data, so as to achieve the purpose of receiving and transmitting the multi-path online classroom audio data by the adaptive network mixing server, thereby realizing the technical effect of improving the smoothness and quality of the online classroom audio data transmission, and further solving the technical problem that the audio data processing method provided by the related art only processes the low-band audio in the normal network state, resulting in low transmission efficiency and poor transmission quality of the audio data.
[0149] In summary, the embodiment of the application focuses on improving the smoothness of online classroom audio transmission in a weak network environment based on the characteristics of the ANMS model, such as anti-packet loss, low bandwidth, and high latency; and improving the quality of online classroom audio (full-band) transmission in a normal network and high-bandwidth environment based on the mixing function of the ANMS model.
[0150] Therefore, the above method provided by the embodiment of the application solves at least the technical problem that the online classroom audio data transmission smoothness can be improved in the case of high packet loss, low bandwidth, and high latency of the online classroom uplink network state, and the online classroom audio data transmission quality can be improved in the case of normal uplink network state or downlink network state (normal bandwidth or high bandwidth) of the online classroom.
[0151] The technical solution provided by the application has at least the following beneficial effects: the anti-packet loss capability, low-bandwidth transmission capability, and high-latency transmission capability of online classroom audio data transmission in a weak network environment can be improved, thereby improving the smoothness of online classroom audio data transmission in a weak network environment.
[0152] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0153] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, and of course it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk) and includes a plurality of instructions for causing an end device (which can be a mobile phone, a computer, a server, or a network device) to execute the method described in each embodiment of the present application.
[0154] Embodiment 2
[0155] According to the embodiments of the present application, a device embodiment for implementing the above-mentioned audio data processing method is also provided, Figure 7 is a structural schematic diagram of an audio data processing device according to an embodiment of the present application, as Figure 7 shown, the device comprises a receiving module 701, a processing module 702 and a sending module 703, wherein,
[0156] The receiving module 701 is configured to receive multiple first audio data from multiple first terminals, wherein the audio processing parameters of each first audio data are adaptively adjusted based on the uplink network state between each first terminal and the adaptive network mixing server.
[0157] The processing module 702 is configured to perform mixing processing on the multiple first audio data to obtain second audio data.
[0158] The sending module 703 is configured to send the second audio data to multiple second terminals, wherein the audio processing parameters of the second audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second terminal.
[0159] Optionally, in the audio data processing apparatus, the audio data processing further comprises: evaluating, respectively, uplink network states between each of the first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and feeding back the first evaluation results to the first terminals, so that each of the first terminals adaptively adjusts the audio processing parameters of the corresponding first audio data by using the first evaluation results.
[0160] Optionally, in the audio data processing apparatus, the audio data processing further comprises: evaluating, respectively, uplink network states between each of the first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and adaptively adjusting the first network jitter equalization parameters of each of the first terminals based on the first evaluation results.
[0161] Optionally, in the audio data processing apparatus, the audio data processing further comprises: evaluating, respectively, uplink network states between each of the first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and limiting the packet loss compensation data of each of the first terminals based on the first evaluation results.
[0162] Optionally, in the audio data processing apparatus, the audio data processing further comprises: receiving, respectively, second evaluation results from each of the second terminals, wherein the second evaluation results are used for evaluating downlink network states between each of the second terminals and the adaptive network mixing server, and the second evaluation results comprise downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation; and adaptively adjusting the audio processing parameters of the second audio data of each of the second terminals based on the second evaluation results.
[0163] Optionally, in the audio data processing apparatus, the second evaluation results are further used for adaptively adjusting the second network jitter equalization parameters of each of the second terminals, and limiting the packet loss compensation data of each of the second terminals.
[0164] Optionally, the processing module 702 is further configured to: mix the first audio data to obtain third audio data; and perform post-processing corresponding to each of the first terminals on the third audio data to obtain the second audio data, wherein the post-processing comprises at least one of the following: reverberation processing, gain adjustment.
[0165] Optionally, Figure 8is a structural schematic diagram of an optional audio data processing device according to an embodiment of the application, as shown in the figure, the device comprises all the modules shown in the figure Figure 8 Figure 7 In addition to all the modules shown in the figure, the device further comprises a display module 704 configured to: in response to a first touch operation on the graphical user interface, set a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each first terminal to adaptively adjust the audio processing parameter of each first audio data based on the uplink network state, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameter of the second audio data based on the downlink network state; in response to a second touch operation on the graphical user interface, select multiple first audio data from the candidate audio data set; in response to a third touch operation on the graphical user interface, perform mixing processing on the multiple first audio data to obtain the second audio data; and in response to a fourth touch operation on the graphical user interface, send the second audio data to the multiple second terminals.
[0166] It should be noted that the receiving module 701, the processing module 702 and the sending module 703 correspond to steps S21 to S23 in Embodiment 1, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules can run in the computer terminal 10 provided in Embodiment 1 as part of the device.
[0167] In the embodiment of the application, multiple first audio data are received from multiple first terminals, wherein the audio processing parameter of each first audio data is adaptively adjusted based on the uplink network state between each first terminal and the adaptive network mixing server, the second audio data is obtained by mixing processing the multiple first audio data, and the second audio data is further sent to multiple second terminals, wherein the audio processing parameter of the second audio data is adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second terminal.
[0168] It is easy to note that, through the embodiment of the application, the uplink and downlink network states are adaptively adjusted based on the adaptive network mixing server to perform mixing processing and network transmission on the audio data, so as to achieve the purpose of receiving and sending multiple audio data through the adaptive network mixing server, thereby realizing the technical effect of improving the smoothness and quality of audio data transmission, and further solving the technical problem of the audio data processing method provided by the related art, which only processes low-frequency band audio in normal network state, resulting in low transmission efficiency and poor transmission quality of audio data.
[0169] According to the embodiment of the application, a device embodiment for implementing the above-mentioned audio data processing method is further provided,Figure 9 is a structural schematic diagram of another audio data processing apparatus according to an embodiment of the present application, as shown in Figure 9 the apparatus comprises a receiving module 901, a processing module 902 and a sending module 903, wherein,
[0170] the receiving module 901 is configured to receive multiple pieces of first online conference audio data from multiple first online conference terminals, wherein the audio processing parameters of each piece of first online conference audio data are adaptively adjusted based on the uplink network status between each first online conference terminal and the adaptive network mixing server;
[0171] the processing module 902 is configured to perform mixing processing on the multiple pieces of first online conference audio data to obtain second online conference audio data;
[0172] the sending module 903 is configured to send the second online conference audio data to multiple second online conference terminals, wherein the audio processing parameters of the second online conference audio data are adaptively adjusted based on the downlink network status between the adaptive network mixing server and each second online conference terminal.
[0173] Optionally, Figure 10 is a structural schematic diagram of another optional audio data processing apparatus according to an embodiment of the present application, as shown in Figure 10 the apparatus comprises all the modules shown in Figure 9 and further comprises a display module 904 configured to, in response to a first touch operation on a graphical user interface, set a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each first online conference terminal to adaptively adjust the audio processing parameters of each piece of first online conference audio data based on the uplink network status, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameters of the second online conference audio data based on the downlink network status; in response to a second touch operation on the graphical user interface, select multiple pieces of first online conference audio data from the candidate online conference audio data set; in response to a third touch operation on the graphical user interface, perform mixing processing on the multiple pieces of first online conference audio data to obtain second online conference audio data; and in response to a fourth touch operation on the graphical user interface, send the second online conference audio data to multiple second online conference terminals.
[0174] It should be noted that the receiving module 901, the processing module 902 and the sending module 903 correspond to steps S51 to S53 in Embodiment 1, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules can run in the computer terminal 10 provided in Embodiment 1 as part of the apparatus.
[0175] In the embodiment of the present application, multi-channel first online conference audio data from a plurality of first online conference terminals is received, wherein the audio processing parameters of each channel of the first online conference audio data are adaptively adjusted based on the uplink network status between each first online conference terminal and the adaptive network mixing server, the second online conference audio data is obtained by mixing the multi-channel first online conference audio data, and the second online conference audio data is further transmitted to a plurality of second online conference terminals, wherein the audio processing parameters of the second online conference audio data are adaptively adjusted based on the downlink network status between the adaptive network mixing server and each second online conference terminal.
[0176] It is easy to note that, through the embodiment of the present application, the uplink and downlink network status during the online conference is adaptively adjusted based on the adaptive network mixing server to mix and transmit the online conference audio data, so as to achieve the purpose of receiving and transmitting multi-channel online conference audio data through the adaptive network mixing server, thereby realizing the technical effect of improving the smoothness and quality of online conference audio data transmission, and further solving the technical problem that the audio data processing method provided by the related art only processes low-frequency band audio in normal network status, resulting in low transmission efficiency and poor transmission quality.
[0177] According to the embodiment of the present application, an apparatus embodiment for implementing the above-mentioned audio data processing method is also provided, Figure 11 is a structural schematic diagram of another audio data processing apparatus according to the embodiment of the present application, as Figure 11 shown, the apparatus comprises a receiving module 1101, a processing module 1102 and a sending module 1103, wherein,
[0178] The receiving module 1101 is configured to receive multi-channel first online classroom audio data from a plurality of first online classroom terminals, wherein the audio processing parameters of each channel of the first online classroom audio data are adaptively adjusted based on the uplink network status between each first online classroom terminal and the adaptive network mixing server.
[0179] The processing module 1102 is configured to mix the multi-channel first online classroom audio data to obtain second online classroom audio data.
[0180] The sending module 1103 is configured to transmit the second online classroom audio data to a plurality of second online classroom terminals, wherein the audio processing parameters of the second online classroom audio data are adaptively adjusted based on the downlink network status between the adaptive network mixing server and each second online classroom terminal.
[0181] It should be noted that the receiving module 1101, the processing module 1102 and the sending module 1103 correspond to steps S61 to S63 in Embodiment 1, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the disclosure of the above embodiment. It should be noted that the above modules can run in the computer terminal 10 provided in Embodiment 1 as part of the device.
[0182] In the embodiment of the present application, multiple first online classroom audio data from multiple first online classroom terminals are received, wherein the audio processing parameters of each first online classroom audio data are adaptively adjusted based on the uplink network state between each first online classroom terminal and the adaptive network audio mixing server. The second online classroom audio data is obtained by mixing the multiple first online classroom audio data, and the second online classroom audio data is further sent to multiple second online classroom terminals, wherein the audio processing parameters of the second online classroom audio data are adaptively adjusted based on the downlink network state between the adaptive network audio mixing server and each second online classroom terminal.
[0183] It is easy to note that, through the embodiment of the present application, the uplink and downlink network states in the online classroom process are adaptively adjusted based on the adaptive network audio mixing server to mix and transmit the online classroom audio data, so as to achieve the purpose of receiving and sending multiple online classroom audio data through the adaptive network audio mixing server, thereby realizing the technical effect of improving the smoothness and quality of online classroom audio data transmission, and further solving the technical problem that the audio data processing method provided by the related art only processes low-frequency band audio in normal network state, resulting in low audio data transmission efficiency and poor transmission quality.
[0184] It should be noted that the preferred embodiments of the present embodiment can refer to the related description in Embodiment 1, which will not be repeated here.
[0185] Embodiment 3
[0186] According to the embodiments of the present application, an embodiment of an electronic device is also provided, which can be any one of the computing devices in the computing device group. The electronic device comprises a processor and a memory, wherein:
[0187] The memory is connected with the processor and is configured to provide the processor with instructions for processing the following steps: receiving multiple pieces of first audio data from multiple first terminals, wherein audio processing parameters of each piece of first audio data are adaptively adjusted based on uplink network states between each first terminal and the adaptive network mixing server; mixing the multiple pieces of first audio data to obtain second audio data; and sending the second audio data to multiple second terminals, wherein audio processing parameters of the second audio data are adaptively adjusted based on downlink network states between the adaptive network mixing server and each second terminal.
[0188] In the embodiment of the present application, multiple pieces of first audio data are received from multiple first terminals, wherein audio processing parameters of each piece of first audio data are adaptively adjusted based on uplink network states between each first terminal and the adaptive network mixing server, the multiple pieces of first audio data are mixed to obtain second audio data, and the second audio data is further sent to multiple second terminals, wherein audio processing parameters of the second audio data are adaptively adjusted based on downlink network states between the adaptive network mixing server and each second terminal.
[0189] It is easy to note that, by the embodiment of the present application, the uplink and downlink network states are adaptively adjusted based on the adaptive network mixing server to mix and transmit the audio data, the purpose of receiving and sending multiple pieces of audio data by the adaptive network mixing server is achieved, the technical effect of improving the smoothness and quality of audio data transmission is achieved, and the technical problem of the audio data processing method provided by the related art only processing low-frequency band audio in a normal network state, resulting in low transmission efficiency and poor transmission quality of audio data is solved.
[0190] It should be noted that the preferred embodiments of the present embodiment can refer to the related description in Embodiment 1, which will not be repeated here.
[0191] Embodiment 4
[0192] The embodiment of the present application can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Alternatively, in the present embodiment, the computer terminal can be replaced by a mobile terminal or other terminal device.
[0193] Alternatively, in the present embodiment, the computer terminal can be located in at least one network device of multiple network devices of a computer network.
[0194] In the embodiment, the computer terminal can execute program codes of the following steps in the audio data processing method: receiving multiple pieces of first audio data from multiple first terminals, wherein audio processing parameters of each piece of first audio data are adaptively adjusted based on uplink network states between each first terminal and the adaptive network mixing server; performing mixing processing on the multiple pieces of first audio data to obtain second audio data; and sending the second audio data to multiple second terminals, wherein audio processing parameters of the second audio data are adaptively adjusted based on downlink network states between the adaptive network mixing server and each second terminal.
[0195] Optionally, Figure 12 is a structural block diagram of another computer terminal according to an embodiment of the present application, as shown in the figure, the computer terminal can include one or more (only one is shown in the figure) processors 122, a memory 124, and a peripheral interface 126. Figure 12
[0196] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the audio data processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned audio data processing method. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the computer terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0197] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: receiving multiple pieces of first audio data from multiple first terminals, wherein audio processing parameters of each piece of first audio data are adaptively adjusted based on uplink network states between each first terminal and the adaptive network mixing server; performing mixing processing on the multiple pieces of first audio data to obtain second audio data; and sending the second audio data to multiple second terminals, wherein audio processing parameters of the second audio data are adaptively adjusted based on downlink network states between the adaptive network mixing server and each second terminal.
[0198] Optionally, the processor can further execute program codes of the following steps: respectively evaluating uplink network states between each of the first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and feeding back the first evaluation results to the first terminals, so that each of the first terminals adaptively adjusts audio processing parameters of the corresponding first audio data by using the first evaluation results.
[0199] Optionally, the processor can further execute program codes of the following steps: respectively evaluating uplink network states between each of the first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and adaptively adjusting the first network jitter equalization parameters of each of the first terminals based on the first evaluation results.
[0200] Optionally, the processor can further execute program codes of the following steps: respectively evaluating uplink network states between each of the first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and limiting the packet loss compensation data of each of the first terminals based on the first evaluation results.
[0201] Optionally, the processor can further execute program codes of the following steps: respectively receiving second evaluation results from each of the second terminals, wherein the second evaluation results are used for evaluating downlink network states between each of the second terminals and the adaptive network mixing server, and the second evaluation results comprise downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation; and adaptively adjusting audio processing parameters of the second audio data of each of the second terminals based on the second evaluation results.
[0202] Optionally, the processor can further execute program codes of the following steps: the second evaluation results are also used for adaptively adjusting the second network jitter equalization parameters of each of the second terminals, and limiting the packet loss compensation data of each of the second terminals.
[0203] Optionally, the processor can further execute program codes of the following steps: mixing the first audio data to obtain third audio data; and performing post-processing corresponding to each of the first terminals on the third audio data to obtain the second audio data, wherein the post-processing comprises at least one of the following: reverberation processing, gain adjustment.
[0204] Optionally, the processor can further execute program codes of the following steps: in response to a first touch operation on the graphical user interface, setting a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each first terminal to adaptively adjust the audio processing parameter of each first audio data based on the uplink network state, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameter of the second audio data based on the downlink network state; in response to a second touch operation on the graphical user interface, selecting multiple first audio data from the candidate audio data set; in response to a third touch operation on the graphical user interface, mixing the multiple first audio data to obtain the second audio data; and in response to a fourth touch operation on the graphical user interface, sending the second audio data to the multiple second terminals.
[0205] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: receiving multiple first online conference audio data from multiple first online conference terminals, wherein the audio processing parameter of each first online conference audio data is adaptively adjusted based on the uplink network state between each first online conference terminal and the adaptive network mixing server; mixing the multiple first online conference audio data to obtain second online conference audio data; and sending the second online conference audio data to multiple second online conference terminals, wherein the audio processing parameter of the second online conference audio data is adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second online conference terminal.
[0206] Optionally, the processor can further execute program codes of the following steps: in response to a first touch operation on the graphical user interface, setting a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each first terminal to adaptively adjust the audio processing parameter of each first audio data based on the uplink network state, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameter of the second audio data based on the downlink network state; in response to a second touch operation on the graphical user interface, selecting multiple first audio data from the candidate audio data set; in response to a third touch operation on the graphical user interface, mixing the multiple first audio data to obtain the second audio data; and in response to a fourth touch operation on the graphical user interface, sending the second audio data to the multiple second terminals.
[0207] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: receiving multiple pieces of first online class audio data from multiple first online class terminals, wherein audio processing parameters of each piece of first online class audio data are adaptively adjusted based on uplink network states between each first online class terminal and the adaptive network audio mixing server; mixing and processing the multiple pieces of first online class audio data to obtain second online class audio data; and sending the second online class audio data to multiple second online class terminals, wherein audio processing parameters of the second online class audio data are adaptively adjusted based on downlink network states between the adaptive network audio mixing server and each second online class terminal.
[0208] In the embodiment of the present application, multiple pieces of first audio data are received from multiple first terminals, wherein audio processing parameters of each piece of first audio data are adaptively adjusted based on uplink network states between each first terminal and the adaptive network audio mixing server, second audio data is obtained by mixing and processing the multiple pieces of first audio data, and the second audio data is further sent to multiple second terminals, wherein audio processing parameters of the second audio data are adaptively adjusted based on downlink network states between the adaptive network audio mixing server and each second terminal.
[0209] It is easy to note that, by the embodiment of the present application, the uplink and downlink network states are adaptively adjusted based on the adaptive network audio mixing server to mix and process the audio data and to transmit the audio data, the purpose of receiving and sending multiple pieces of audio data by the adaptive network audio mixing server is achieved, the technical effect of improving the smoothness and quality of audio data transmission is achieved, and the technical problem of the audio data processing method provided by the related art that only processes low-frequency band audio in a normal network state to cause low audio data transmission efficiency and poor transmission quality is solved.
[0210] Those skilled in the art can understand that, Figure 12 The structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or the like. Figure 12 It does not limit the structure of the electronic device. For example, the computer terminal can further include more or fewer components (such as a network interface, a display device, etc.) than Figure 12 shown, or have a different configuration than Figure 12 shown.
[0211] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by instructing the terminal device related hardware through programs, and the programs can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0212] According to the embodiments of the present application, an embodiment of a computer readable storage medium is also provided. Optionally, in the present embodiment, the computer readable storage medium can be used to store the program code executed by the audio data processing method provided in the above-mentioned embodiment 1.
[0213] Optionally, in the present embodiment, the computer readable storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0214] Optionally, in the present embodiment, the computer readable storage medium is configured to store program code for performing the following steps: receiving a plurality of first audio data from a plurality of first terminals, wherein the audio processing parameters of each of the first audio data are adaptively adjusted based on the uplink network status between each of the first terminals and the adaptive network mixing server; mixing the plurality of first audio data to obtain second audio data; and sending the second audio data to a plurality of second terminals, wherein the audio processing parameters of the second audio data are adaptively adjusted based on the downlink network status between the adaptive network mixing server and each of the second terminals.
[0215] Optionally, in the present embodiment, the computer readable storage medium is configured to store program code for performing the following steps: respectively evaluating the uplink network status between each of the first terminals and the adaptive network mixing server to obtain a first evaluation result, wherein the first evaluation result includes: uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and feeding back the first evaluation result to the plurality of first terminals, so that each of the plurality of first terminals adaptively adjusts the audio processing parameters of the corresponding first audio data using the first evaluation result.
[0216] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: respectively evaluating uplink network status between each of the first terminals and the adaptive network mixing server to obtain a first evaluation result, wherein the first evaluation result comprises uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and adaptively adjusting a first network jitter equalization parameter corresponding to each of the first terminals based on the first evaluation result.
[0217] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: respectively evaluating uplink network status between each of the first terminals and the adaptive network mixing server to obtain a first evaluation result, wherein the first evaluation result comprises uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; and limiting a packet loss compensation data corresponding to each of the first terminals based on the first evaluation result.
[0218] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: respectively receiving a second evaluation result from each of the second terminals, wherein the second evaluation result is used for evaluating downlink network status between each of the second terminals and the adaptive network mixing server, and the second evaluation result comprises downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation; and adaptively adjusting an audio processing parameter of second audio data corresponding to each of the second terminals based on the second evaluation result.
[0219] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: the second evaluation result is further used for adaptively adjusting a second network jitter equalization parameter by each of the second terminals, and limiting a packet loss compensation data by each of the second terminals.
[0220] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: mixing the first audio data to obtain third audio data; and performing post-processing corresponding to each of the first terminals on the third audio data to obtain the second audio data, wherein the post-processing comprises at least one of the following: reverb processing, gain adjustment.
[0221] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: in response to a first touch operation on the graphical user interface, setting a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each first terminal to adaptively adjust the audio processing parameter of each first audio data based on the uplink network state, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameter of the second audio data based on the downlink network state; in response to a second touch operation on the graphical user interface, selecting multiple first audio data from the candidate audio data set; in response to a third touch operation on the graphical user interface, mixing the multiple first audio data to obtain the second audio data; and in response to a fourth touch operation on the graphical user interface, sending the second audio data to the multiple second terminals.
[0222] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: receiving multiple first online conference audio data from multiple first online conference terminals, wherein the audio processing parameter of each first online conference audio data is adaptively adjusted based on the uplink network state between each first online conference terminal and the adaptive network mixing server; mixing the multiple first online conference audio data to obtain second online conference audio data; and sending the second online conference audio data to multiple second online conference terminals, wherein the audio processing parameter of the second online conference audio data is adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second online conference terminal.
[0223] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: in response to a first touch operation on the graphical user interface, setting a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each first online conference terminal to adaptively adjust the audio processing parameter of each first online conference audio data based on the uplink network state, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameter of the second online conference audio data based on the downlink network state; in response to a second touch operation on the graphical user interface, selecting multiple first online conference audio data from the candidate online conference audio data set; in response to a third touch operation on the graphical user interface, mixing the multiple first online conference audio data to obtain the second online conference audio data; and in response to a fourth touch operation on the graphical user interface, sending the second online conference audio data to the multiple second online conference terminals.
[0224] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: receiving multiple pieces of first online class audio data from multiple first online class terminals, wherein audio processing parameters of each piece of first online class audio data are adaptively adjusted based on uplink network status between each first online class terminal and the adaptive network mixing server; mixing the multiple pieces of first online class audio data to obtain second online class audio data; and sending the second online class audio data to multiple second online class terminals, wherein audio processing parameters of the second online class audio data are adaptively adjusted based on downlink network status between the adaptive network mixing server and each second online class terminal.
[0225] The serial numbers of the above embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0226] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0227] In the several embodiments of the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the above-described device embodiments are only illustrative, and the division of units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.
[0228] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiment.
[0229] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0230] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0231] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. An audio data processing method, characterized by, The method comprises: receiving multiple pieces of first audio data from multiple first terminals, wherein audio processing parameters of each piece of first audio data are adaptively adjusted based on uplink network status between each first terminal and an adaptive network mixing server, comprising: respectively evaluating uplink network status between each first terminal in the multiple first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; feeding back the first evaluation results to the multiple first terminals, so that each first terminal in the multiple first terminals adaptively adjusts audio processing parameters of corresponding first audio data using the first evaluation results; performing mixing processing on the multiple pieces of first audio data to obtain second audio data; sending the second audio data to multiple second terminals, wherein audio processing parameters of the second audio data are adaptively adjusted based on downlink network status between the adaptive network mixing server and each second terminal, comprising: respectively receiving second evaluation results from each second terminal in the multiple second terminals, wherein the second evaluation results are used for evaluating downlink network status between each second terminal and the adaptive network mixing server, and the second evaluation results comprise downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation; adaptively adjusting audio processing parameters of corresponding second audio data of each second terminal based on the second evaluation results.
2. The audio data processing method of claim 1, wherein, The audio data processing further comprises: respectively evaluating uplink network status between each first terminal in the multiple first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; adaptively adjusting first network jitter equalization parameters of each first terminal in the multiple first terminals based on the first evaluation results.
3. The audio data processing method of claim 1, wherein, The audio data processing further comprises: respectively evaluating uplink network status between each first terminal in the multiple first terminals and the adaptive network mixing server to obtain first evaluation results, wherein the first evaluation results comprise uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; throttling packet loss compensation data of each first terminal in the multiple first terminals based on the first evaluation results.
4. The audio data processing method of claim 1, wherein, The second evaluation results are further used for adaptively adjusting second network jitter equalization parameters of each second terminal in the multiple second terminals, and throttling packet loss compensation data of each second terminal in the multiple second terminals.
5. The audio data processing method of claim 1, wherein, The mixing processing on the multiple pieces of first audio data to obtain the second audio data comprises: performing mixing processing on the multiple pieces of first audio data to obtain third audio data; performing post-processing corresponding to each first terminal on the third audio data to obtain the second audio data, wherein the post-processing comprises at least one of the following: reverb processing, gain adjustment.
6. The audio data processing method of claim 1, wherein, An adaptive network mixing server provides a graphical user interface, which displays at least in part an audio processing scene, the audio data processing method further comprises: in response to a first touch operation on the graphical user interface, setting a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each first terminal to adaptively adjust the audio processing parameters of each first audio data based on the uplink network state, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameters of the second audio data based on the downlink network state; in response to a second touch operation on the graphical user interface, selecting the multi-channel first audio data from the candidate audio data set; in response to a third touch operation on the graphical user interface, mixing the multi-channel first audio data to obtain the second audio data; in response to a fourth touch operation on the graphical user interface, sending the second audio data to the plurality of second terminals.
7. An audio data processing method, characterized by, comprises: receiving multi-channel first online conference audio data from a plurality of first online conference terminals, wherein the audio processing parameters of each channel of first online conference audio data are adaptively adjusted based on the uplink network state between each first online conference terminal and the adaptive network mixing server, including: respectively evaluating the uplink network state between each first online conference terminal in the plurality of first online conference terminals and the adaptive network mixing server to obtain a first evaluation result, wherein the first evaluation result includes: uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; feeding back the first evaluation result to the plurality of first online conference terminals, so that each first online conference terminal in the plurality of first online conference terminals adaptively adjusts the audio processing parameters of the corresponding first online conference audio data using the first evaluation result; mixing the multi-channel first online conference audio data to obtain second online conference audio data; sending the second online conference audio data to a plurality of second online conference terminals, wherein the audio processing parameters of the second online conference audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second online conference terminal, including: respectively receiving a second evaluation result from each second online conference terminal in the plurality of second online conference terminals, wherein the second evaluation result is used to evaluate the downlink network state between each second online conference terminal and the adaptive network mixing server, and the second evaluation result includes: downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation; adaptively adjusting the audio processing parameters of the corresponding second online conference audio data of each second online conference terminal based on the second evaluation result.
8. The audio data processing method of claim 7, wherein, A graphical user interface is provided by an adaptive network mixing server, the graphical user interface displays at least partially an online conference audio processing scene, the audio data processing method further comprises: in response to a first touch operation on the graphical user interface, setting a first parameter adjustment option and a second parameter adjustment option, wherein the first parameter adjustment option is used to trigger each first online conference terminal to adaptively adjust the audio processing parameters of each first online conference audio data based on the uplink network state, and the second parameter adjustment option is used to trigger the adaptive network mixing server to adaptively adjust the audio processing parameters of the second online conference audio data based on the downlink network state; in response to a second touch operation on the graphical user interface, selecting the multiple first online conference audio data from a candidate online conference audio data set; in response to a third touch operation on the graphical user interface, mixing the multiple first online conference audio data to obtain the second online conference audio data; in response to a fourth touch operation on the graphical user interface, sending the second online conference audio data to the multiple second online conference terminals.
9. An audio data processing method, characterized by, comprises: receiving multiple first online classroom audio data from multiple first online classroom terminals, wherein the audio processing parameters of each first online classroom audio data are adaptively adjusted based on the uplink network state between each first online classroom terminal and the adaptive network mixing server, including: respectively evaluating the uplink network state between each first online classroom terminal in the multiple first online classroom terminals and the adaptive network mixing server to obtain a first evaluation result, wherein the first evaluation result includes: uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; feeding back the first evaluation result to the multiple first online classroom terminals, so that each first online classroom terminal in the multiple first online classroom terminals adaptively adjusts the audio processing parameters of the corresponding first online classroom audio data using the first evaluation result; mixing the multiple first online classroom audio data to obtain second online classroom audio data; sending the second online classroom audio data to multiple second online classroom terminals, wherein the audio processing parameters of the second online classroom audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second online classroom terminal, including: respectively receiving a second evaluation result from each second online classroom terminal in the multiple second online classroom terminals, wherein the second evaluation result is used to evaluate the downlink network state between each second online classroom terminal and the adaptive network mixing server, and the second evaluation result includes: downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation; adaptively adjusting the audio processing parameters of the corresponding second online classroom audio data of each second online classroom terminal based on the second evaluation result.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored program, wherein the program controls a device where the computer readable storage medium is located to perform the audio data processing method in any one of claims 1 to 9 when the program is running.
11. An electronic device, comprising: Comprise: a processor; and a memory connected with the processor, for providing the processor with instructions for processing the following processing steps: receiving multiple pieces of first audio data from multiple first terminals, wherein the audio processing parameters of each piece of first audio data are adaptively adjusted based on the uplink network state between each first terminal and the adaptive network mixing server, comprising: respectively evaluating the uplink network state between each first terminal in the multiple first terminals and the adaptive network mixing server to obtain a first evaluation result, wherein the first evaluation result comprises: uplink network bandwidth evaluation, uplink network packet loss evaluation and uplink network delay evaluation; feeding back the first evaluation result to the multiple first terminals, so that each first terminal in the multiple first terminals adaptively adjusts the audio processing parameters of the corresponding first audio data using the first evaluation result; mixing the multiple pieces of first audio data to obtain second audio data; sending the second audio data to multiple second terminals, wherein the audio processing parameters of the second audio data are adaptively adjusted based on the downlink network state between the adaptive network mixing server and each second terminal, comprising: respectively receiving a second evaluation result from each second terminal in the multiple second terminals, wherein the second evaluation result is used to evaluate the downlink network state between each second terminal and the adaptive network mixing server, and the second evaluation result comprises: downlink network bandwidth evaluation, downlink network packet loss evaluation and downlink network delay evaluation; adaptively adjusting the audio processing parameters of the corresponding second audio data of each second terminal based on the second evaluation result.
Citation Information
Patent Citations
Multimedia conference control method and server
CN105812717A
Voice coding control method and device, and storage medium
CN111951813A