Data processing method and related equipment
By integrating audio delay information between the first device and the second device, the universality and accuracy issues of end-to-end delay statistics in the existing technology are solved, and detailed and comprehensive statistics of end-to-end delay in a real usage environment are achieved, thereby improving the accuracy and coverage of delay data.
Patent Information
- Application Number
- CN202410283881.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies have limited versatility and accuracy when calculating end-to-end delays, making it difficult to accurately obtain delay data across different devices and in large-scale real-time communication scenarios.
By transmitting audio data packets between the first device and the second device, the delay information on each device is counted and integrated into the target audio delay. This method has high accuracy and is suitable for any device and large-scale real-time communication scenarios.
It achieves detailed and comprehensive statistics of end-to-end delay in real-world usage environments, improves the accuracy and coverage of delay data, and is suitable for various real-time communication scenarios, including audio calls, video calls, live broadcasts, online conferences, and choruses.
Smart Images

Figure CN120658655A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method and related equipment. Background Art
[0002] With the development of the internet, real-time communication scenarios built on the internet in our daily lives have greatly facilitated information transmission and improved efficiency. For example, in real-time communication scenarios such as audio and video calls, interactive live broadcasts, and choral singing (i.e., communication scenarios with real-time requirements), online real-time audio and video interactions can be carried out. In such real-time communication scenarios, audio signals are transmitted between different devices and there is end-to-end latency. End-to-end latency is an important and critical metric that can guide long-term improvements in real-time communication scenarios. If the latency is too large, the real-time performance of the real-time communication scenario will be reduced, affecting the service experience in some interactive scenarios. Therefore, calculating end-to-end latency is a very important task. The current methods used in the industry to calculate end-to-end latency are limited in both versatility and accuracy. Summary of the Invention
[0003] The embodiments of the present application provide a data processing method and related equipment, which are not limited by usage scenarios and usage scales, have high versatility, and can improve the accuracy and comprehensiveness of delay data.
[0004] In one aspect, an embodiment of the present application provides a data processing method, the method comprising:
[0005] During real-time communication between a first device and a second device, receiving an audio data packet transmitted by the first device; the audio data packet includes an audio signal on the first device and first delay information, the first delay information being used to indicate delays incurred at various transmission processing stages on the first device during transmission of the audio signal along the audio link;
[0006] Performing delay statistical processing on the audio data packet to obtain second delay information of the audio signal; the second delay information is used to indicate: delays incurred at various receiving and processing stages experienced by the audio signal on the second device during transmission along the audio link;
[0007] The first delay information of the audio signal and the second delay information of the audio signal are integrated to obtain a target audio delay.
[0008] In one aspect, an embodiment of the present application provides a data processing device, comprising:
[0009] A receiving unit, configured to receive an audio data packet transmitted by the first device during real-time communication between the first device and the second device; the audio data packet including an audio signal from the first device and first delay information, the first delay information being used to indicate delays incurred at various transmission processing stages experienced by the first device during transmission of the audio signal along the audio link;
[0010] a processing unit, configured to perform delay statistical processing on the audio data packet to obtain second delay information of the audio signal; the second delay information being used to indicate delays incurred at various receiving and processing stages experienced by the second device during transmission of the audio signal along the audio link;
[0011] The processing unit is further configured to integrate the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay.
[0012] In one aspect, an embodiment of the present application provides a computer device, comprising:
[0013] a processor suitable for executing a computer program;
[0014] Computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the above-mentioned data processing method is implemented.
[0015] On the one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is loaded by a processor and executes the above-mentioned data processing method.
[0016] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the above-mentioned data processing method is implemented.
[0017] In an embodiment of the present application, the end-to-end delay of the audio signal transmitted between any two devices can be counted during the real-time communication process. Specifically, based on the delay generated by the audio signal in the processing stages experienced by the audio signal on different device sides during the transmission of the entire audio link, the delay generated by the audio signal in all processing stages of the audio link can be counted. By integrating the delay generated by the audio signal in all processing stages of the audio link, the final target audio delay can be obtained. In this way, a detailed and comprehensive end-to-end delay can be obtained, and the accuracy of the audio delay can be improved. For any two devices in the same real-time communication scenario, the above scheme can be used to obtain the corresponding end-to-end delay. That is, no matter how large the real-time communication scenario is, this scheme can cover all devices in the real-time communication scenario for delay statistics, so it is not limited by the scale of use. In addition, no matter how large the differences in device characteristics are, accurate delays can be obtained in real-time usage scenarios. The usage scenarios are not limited to artificially constructed test scenarios. Based on the unlimited usage scale and usage scenarios, it has high versatility and high statistical efficiency for large-scale real-time communication scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1a is a schematic diagram of an audio link provided by an exemplary embodiment of the present application;
[0020] Figure 1b is an architectural diagram of a data processing system provided by an exemplary embodiment of the present application;
[0021] Figure 2 is a flowchart of a data processing method provided by an exemplary embodiment of the present application;
[0022] Figure 3 This is a schematic diagram of a data processing scenario provided by an exemplary embodiment of the present application;
[0023] Figure 4 is a flowchart of another data processing method provided by an exemplary embodiment of the present application;
[0024] Figure 5a This is a schematic diagram of an audio acquisition processing flow provided by an exemplary embodiment of the present application;
[0025] Figure 5bThis is a schematic diagram of a device playback processing flow provided by an exemplary embodiment of the present application;
[0026] Figure 6 is a schematic diagram of an audio link delay provided by an exemplary embodiment of the present application;
[0027] Figure 7 is a structural diagram of a data processing device provided by an exemplary embodiment of the present application;
[0028] Figure 8 It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] In this application, the terms "first," "second," and so on are used to distinguish identical or similar items with substantially the same purpose or function. It should be understood that "first," "second," and "nth" do not have a logical or temporal dependency, nor do they limit quantity or order of execution. The term "at least one" means one or more, and "plurality" means two or more; for example, "at least one audio signal" refers to one, two, or more audio signals.
[0031] In this application, the term "module" or "unit" refers to a computer program or part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the functions of the module or unit.
[0032] The present application proposes a data processing scheme, which relates to a data processing system, method and related equipment. In this scheme, during real-time communication between a first device and a second device, the second device can receive an audio data packet transmitted by the first device. The audio data packet includes an audio signal and first delay information on the first device side. The first delay information is used to indicate the delay data generated by each transmission processing stage experienced by the audio signal on the first device side during transmission according to the audio link. It can be seen that the first delay information obtained by statistics on the first device side can fully indicate the delay of each stage experienced by the first device side during the transmission of the audio signal, which is conducive to the second device obtaining comprehensive and accurate delay data. Then, the second device can perform delay statistics processing on the audio data packet to obtain second delay information. The second delay information is used for the delay data generated by each receiving processing stage experienced by the second device side during transmission according to the audio link. It can be seen that the second delay information obtained by statistics on the second device side can fully indicate the delay of each stage experienced by the second device side during the transmission of the audio signal. Therefore, the target audio delay obtained by integrating the first delay information and the second delay information is a more accurate and detailed end-to-end delay.
[0033] According to the above logic, in the process of transmitting audio signals between two devices, the delay of the audio signal at each stage of the audio link can be accurately measured based on the processing of the audio signal, thereby improving the detail and accuracy of the end-to-end delay. In addition, the data processing solution provided by this application can be directly deployed in an online environment, which is a manual use environment built based on the real use environment, such as the real use environment for an application in the running process. By obtaining the end-to-end delay in the real use environment, the data collection accuracy is high and can cover large-scale online scenarios.
[0034] Real-time communication refers to a communication method that can transmit audio, video, and other information in a short period of time. This communication method requires a low-latency network to ensure instant information transmission and feedback. Real-time communication between any two devices can be achieved through dial-up or through the real-time communication capabilities of the target application (applications with real-time communication capabilities, such as gaming, social networking, or office applications). Business scenarios that support real-time communication can be called real-time communication scenarios, including but not limited to audio calls, video calls, live broadcasts, online meetings, gaming, and chorus scenarios.
[0035] The first device and the second device are both computer devices, specifically terminal devices. The first and second devices can be of different types, for example, the first device is a mobile terminal and the second device is a fixed terminal, or the first device is a tablet computer and the second device is a laptop computer. The first and second devices are any two devices in the same real-time communication scenario; for example, if the real-time communication scenario is a multi-person online conference scenario, then the first and second devices can be any two devices used by participants in the same conference. In a real-time communication scenario, the first device can collect its own audio signal and transmit it to the second device, and the second device can also collect its own audio signal and transmit it to the first device. For example, in a voice call between two devices (such as device A and device B), when the first device is device A (such as a smartphone) and the second device is device B (such as a tablet computer), the sound on device A (including human voice and background sound, etc.) can be transmitted to device B and played out; when the first device is device B and the second device is device A, the sound on device B can also be transmitted to device A and played out. For ease of explanation, this application is explained from the perspective of the second device receiving the audio signal from the first device side, that is, the second device is the executor of the data processing method provided in this application. It can be understood that during the real-time communication between the two devices, the first device can also receive the audio signal from the second device side, and the data processing method provided in this application is also applicable by swapping the first device and the second device.
[0036] The so-called audio link refers to the entire transmission path of the audio signal from its generation to its final playback. In this application, the audio signal involved in the real-time communication between two devices can be transmitted in real time according to the audio link. The audio link includes multiple processing stages, and the processing stage can also be understood as a processing link or a transmission node. There is a serial order between the various processing stages included in the audio link. Based on this serial order, the output of one processing stage can be used as the input of the next processing stage; in each processing stage, the audio signal can be processed directly or based on the result obtained in the previous processing stage. In this application, the device end used to collect audio signals can be called the far end (or the transmitting end, such as the first device), and the device end used to receive audio signals can be called the near end (or the receiving end, such as the second device); the audio end-to-end delay refers to the time delay experienced during the entire process of the audio signal being transmitted from the far end to the near end for playback. In layman's terms, it is the length of time that the audio signal lags behind the far end playback at the near end.
[0037] For ease of understanding, the following is provided Figure 1a The schematic diagram of the audio chain shown in FIG1 illustrates the functions of each processing stage included in the audio chain. Figure 1aAs shown, in an audio link, the far-end processing stages include audio acquisition, preprocessing, encoding, and uplink. Specifically, the far-end device can collect audio data during the audio acquisition stage. Then, during the preprocessing stage, the audio data can be preprocessed. Optionally, this preprocessing can include 3A processing, which collectively refers to several core audio processing technologies, including acoustic echo cancellation (AEC), automatic gain control (AGC), and automatic noise suppression (ANS). Preprocessing audio data can improve audio quality and communication, playing a vital role in various scenarios requiring high-quality audio. After preprocessing, the audio data enters the encoding stage, where it is encoded and then uplinked to the server. The near-end device can pull the audio stream from the server, allowing the audio stream to reach the near-end device. The device on the proximal side can decode the downstream audio stream. In addition, if the device on the proximal side pulls multiple audio streams, it can also decode each audio stream, and then mix the decoded audio data to obtain mixed data, and then send the mixed data to the player for playback. Among them, the so-called audio stream refers to a continuous sequence of audio data. The basic transmission unit of the audio stream can be an audio data packet, which contains audio sampling data within a certain period of time (for example, 20ms or 125ms). In the above links, the audio link includes audio acquisition, preprocessing, encoding, uplink and downlink, decoding, mixing and playback, and these processing stages have their own delays. In this application, the delay of each processing stage in the audio link can be counted, so as to accurately obtain the end-to-end delay of the audio data, and can cover the delay of each audio stream in the processing scenario of multiple audio streams.
[0038] The data processing method provided in this application involves cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing. In the embodiments of this application, cloud computing may be specifically involved. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". The resources in the "cloud" appear to be infinitely scalable to users and can be obtained at any time, used on demand, expanded at any time, and paid for on a per-use basis. For example, the data processing method provided in this application is used to count the delay of the audio signal received by each device in a large-scale real-time communication scenario. The computing power provided by cloud computing can be used to achieve rapid calculation, and the calculated delay can be stored in a database through a storage service for use. The data processing method provided in this application also involves artificial intelligence (AI). For example, when identifying which links in the audio link have caching operations, a neural network model can be used for identification.
[0039] The architecture of the data processing system provided in the embodiments of the present application will be introduced below with reference to the accompanying drawings.
[0040] See Figure 1b , Figure 1b This is an architectural diagram of a data processing system provided by an exemplary embodiment of the present application. Figure 1b As shown, the data processing system includes at least one terminal device (including terminal device 101a, terminal device 101b, ... terminal device 101n) and server 102; each terminal device can establish a communication connection with server 102 via wired or wireless means. Among them, any terminal device includes but is not limited to: smart phones, tablet computers, smart wearable devices, smart voice interaction devices, smart home appliances, personal computers, vehicle-mounted terminals, smart cameras, virtual reality devices and other devices, and this application does not impose any restrictions on this. This application does not impose any restrictions on the number of terminal devices. The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but is not limited to this. This application does not impose any restrictions on the number of servers.
[0041] Assuming that terminal device 101a, terminal device 101b, ..., terminal device 101n are devices in the same real-time communication scenario, and any two devices support mutual transmission of audio signals collected by the respective terminals, any terminal device can execute the data processing method provided in this application. The following, taking the first device as terminal device 101a and the second device as terminal device 101b as an example, briefly describes the interaction process between the first device, the second device, and the server involved in the data processing method provided in this application:
[0042] (1) The terminal device 101a collects audio signals, performs packet processing based on the audio signals, and obtains audio data packets.
[0043] In a specific implementation, the terminal device 101a can collect the audio signal according to the audio link and perform a series of processing on the audio signal (including analog-to-digital conversion, preprocessing and encoding). In the process of processing the audio signal, the terminal device 101a can perform delay statistics on each processing stage on the first device side in the audio link to obtain first delay information, which is used to indicate: the delay generated by each transmission processing stage experienced by the first device side during the transmission of the audio signal according to the audio link. Afterwards, the terminal device 101a can package the audio signal and the first delay information to obtain an audio data packet, and upload the audio data packet to the server 102. Optionally, the server 102 is used to support real-time communication between any two devices in a real-time communication scenario. If the two devices communicate in real time through a target application, then the server 102 can be the background server of the target application, referred to as an application server. Among them, the target application refers to an application with real-time communication functions, and the real-time communication functions include but are not limited to: voice calls, video calls, live broadcasts, chorus and online conferences and other functions. Target applications include but are not limited to any of the following: audio and video applications, social applications, office applications, and educational applications, etc.
[0044] (2) The terminal device 101b receives the audio data packet transmitted by the terminal device 101a.
[0045] In one implementation, the terminal device 101b can send a stream pull request to the server 102, and the stream pull request includes a target identifier. The target identifier can be a user identifier or other device identifier in the same real-time communication scenario, and here it can be the device identifier of the first device (terminal device 101a), such as the device model; the server 102 can determine the audio data packet associated with the target identifier based on the target identifier in the stream pull request, that is, determine the audio data packet of the uplink of the terminal device 101a, so that the server 102 can send the audio data packet to the terminal device 101b, so that the audio stream on the terminal device 101a side can reach the second device and be played in the terminal device 101b. In other words, the terminal device 101b can receive the audio data packet transmitted by the terminal device 101a through the server 102.
[0046] (3) The terminal device 101b performs delay statistics based on the audio data packet to obtain second delay information of the audio signal.
[0047] In one implementation, the terminal device 101b can perform delay statistics at each receiving and processing stage of the audio link based on the audio data packet, obtain the delay data of the audio signal at each receiving and processing stage, and then integrate these delay data to obtain the second delay information of the audio signal. The second delay information can be used to indicate the delay data generated by each receiving and processing stage experienced by the second device side during the transmission of the audio signal according to the audio link. Furthermore, after obtaining the second delay information, the terminal device can report the second delay information to the server 102, so that the server 102 can obtain the end-to-end audio link delay based on the first delay information and the second delay information, and can also obtain the delay of each processing stage on the audio link based on the first delay information and the second delay information, so as to more conveniently understand the audio link environment.
[0048] (4) The terminal device 101b integrates the first delay information and the second delay information to obtain the target audio delay.
[0049] In one embodiment, the terminal device 101b can combine the delays of each transmission stage indicated by the first delay information and the delays of each receiving processing stage indicated by the second delay information to obtain a delay set as the target audio delay, that is, the target audio delay is a set of delays including each processing stage in the audio link, and the target audio delay can be used to indicate the delays generated by the audio signal in each processing stage in the audio link. In another embodiment, the terminal device 101b can sum the delays of each transmission stage indicated by the first delay information and the delays of each receiving processing stage indicated by the second delay information, and the sum obtained can be used as the target audio delay. That is, the target audio delay can also be the sum of the delays of each processing stage of the audio link, and the target delay information is the end-to-end audio delay.
[0050] Furthermore, the target delay information can be used to provide direction guidance for delay optimization. This is because in real-time communication scenarios, if the end-to-end delay is too large, it will affect the user experience. For example, in a chorus scenario, a delay difference of hundreds of milliseconds can be more clearly reflected in the synchronization of lyrics, and a relatively large delay will seriously affect the chorus experience. Targeted optimization of the processing stages in the audio link based on the target audio delay can reduce end-to-end delay, improve the real-time performance and smoothness of the audio signal, and thus improve the user experience. During targeted optimization, the target delay information can be compared with the preset delay threshold to determine whether to optimize and the processing stage for targeted optimization. For example, if the target delay information is greater than the preset delay threshold, the processing stage with the largest delay in the audio link can be found, and the corresponding processing stage can be optimized to improve the processing speed and efficiency of the audio data, thereby reducing the end-to-end audio delay as much as possible and providing a smoother and more real-time audio experience.
[0051] The data processing method provided by this application can be directly deployed in the online environment of the target application. In this way, when the target application is running, the end-to-end audio delay in the real environment can be directly obtained. In this way, on the one hand, it can reflect the delay of the real use scenario, and on the other hand, it does not rely on manual testing and gets rid of the tedious work of manual testing, making it easy to obtain delay data for all devices. On the other hand, the delay of each link in the end-to-end audio link can ensure that the delay is detailed and comprehensive, which also greatly facilitates the viewing and further optimization of delay data.
[0052] The data processing method provided by this application has the following advantages over the solution of manually using corresponding equipment to build a manual environment (such as a live broadcast environment, a communication environment, etc.) for manual testing to obtain end-to-end delay: 1) It is not limited by the characteristic differences between devices. The delay of all devices can be determined by the method provided by this application, and the usage scenarios and devices used will not be limited. 2) It is not limited by the scale of use. If it is a large-scale online scenario, since the method provided by this application can be directly deployed online, it can cover large-scale online scenarios, which can not only collect more accurate delay data, but also has unlimited usage scale. Compared with directly using RTT delay as end-to-end delay in an online environment, the information in this application is detailed and accurate. This is because the method provided by this application can make detailed statistics on the delay data of the intermediate link, and based on the delay data, it can very conveniently optimize and improve the processing logic of the corresponding stage in the audio link. This takes into account not only the audio processing delay, but also the device processing delay. The delay statistics error is small and the accuracy is high.
[0053] Next, the data processing method provided in the embodiment of the present application is introduced.
[0054] See Figure 2 , is a flow chart of a data processing method provided by an exemplary embodiment of the present application. The data processing method can be executed by a second device, which is specifically a computer device (such as Figure 1b In the terminal device 101b), the data processing method may include the following S201-S203.
[0055] S201 , during real-time communication between a first device and a second device, receiving an audio data packet transmitted by the first device.
[0056] In one embodiment, the real-time communication between the first device and the second device includes, but is not limited to, any of the following: real-time communication in a live broadcast scenario, real-time communication in an audio call scenario, real-time communication in a video call scenario, real-time communication in a chorus scenario, real-time communication in a game scenario, and real-time communication in an online conference scenario, etc. This application does not limit the scenarios involved in real-time communication, and for example, it can also include real-time communication in an online education scenario, etc. The above-mentioned live broadcast scenarios, audio call scenarios, chorus scenarios, online conference scenarios, and game scenarios are all real-time communication scenarios, and in the above-mentioned real-time communication scenarios, the transmission of audio signals is supported. The environment involved in the real-time communication between the first device and the second device is a real usage environment, which refers to the environment in which the business is actually performed. For example, the online environment when an audio application is running is a real usage environment. Compared with a manually constructed manual environment, delay statistics in a real usage environment are not limited by differences in device characteristics, have high versatility, and can obtain more reliable and accurate delay data.
[0057] In one implementation, during real-time communication between a first device and a second device, the second device can pull audio data packets transmitted by the first device from a server. Specifically, the second device can send a pull request to the server and then receive the audio data packets transmitted by the first device from the server. This is because the server can be used to manage the audio data packets transmitted by each device in the real-time communication scenario. The transmission of audio signals between any two devices can be achieved by sending audio data packets from the server.
[0058] The audio data packet includes an audio signal on the first device side and first delay information. In this application, the audio data packet may also be referred to as an audio packet. The audio signal on the first device side is stored in a digital form in the audio data packet. The audio data packet has a certain structural organization, which includes but is not limited to: an audio packet header, a packet type, a packet size, audio data, and a first delay information, etc. This information can ensure that the audio signal is correctly decoded and played on various devices that support audio playback, and can enable the second device to obtain the delay data obtained by statistics on the first device side after receiving the audio data packet, so as to determine the end-to-end delay when playing the audio signal on the second device side in combination with the delay data obtained by statistics on the second device side.
[0059] An audio link can be formed between any two devices communicating in real time, allowing the audio signal collected at one end to be transmitted to the other end for playback after being processed according to the various processing stages in the audio link. The various processing stages that the audio signal undergoes on the first device side (including audio acquisition, preprocessing, and encoding, which are the processing stages on the far side) are called transmission processing stages, while the various processing stages that the audio signal undergoes on the second device side in the form of audio data packets (including downlink, decoding, mixing, and playback, which are the near-end stages) are called reception processing stages.
[0060] The first delay information is used to indicate: the delays caused by the audio signal in the various transmission processing stages experienced by the first device side during the transmission according to the audio link. It can be seen that the delays indicated by the first delay information are all uplink delays counted by the first device. In a feasible manner, the first delay information may include the delays of multiple processing stages experienced by the audio signal on the first device side; for example, the first delay information includes acquisition delay, preprocessing delay and encoding delay, which correspond to the acquisition stage, preprocessing stage and encoding stage, respectively.
[0061] S202: Perform delay statistics processing on the audio data packet to obtain second delay information of the audio signal.
[0062] The second delay information indicates the delays incurred during the various receiving and processing stages of the audio signal on the second device during transmission along the audio link. The delays indicated by the second delay information are all downlink delays and may also include delays in audio signal transmission within the network.
[0063] In a specific implementation, the second device can perform a series of processing (including unpacking, decoding, playing, etc.) on the audio data packet according to the audio link. During the processing, the second device can count the delay of each processing stage in the corresponding processing stage, thereby integrating the delays of each processing stage obtained by statistics to obtain the second delay information. During the integration, the delays of each processing stage can be combined to obtain a delay set, which is determined as the second delay information, so that the second delay information includes the delays of each processing stage. In other words, the second delay information can be a data set containing the delays of multiple receiving processing stages. For example, the second delay information includes network transmission delay, decoding delay, and playback delay.
[0064] S203: Integrate the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay.
[0065] Since the first delay information and the second delay information respectively indicate the delays caused by the processing stages experienced by the same audio signal on different device sides, when the first delay information and the second delay information are integrated, in one feasible way, the second device can directly sum the delays caused by the various processing stages indicated by the first delay information and the delays of the various processing stages indicated by the second delay information, and use the sum obtained as the target audio delay. At this time, the target audio delay is the total delay of the audio signal in the entire audio link. In another feasible way, the second device can combine the delays generated by the various processing stages to obtain a delay set, and use the delay set as the target audio delay, so that the target audio delay includes the delays caused by all processing stages experienced by the audio signal during transmission according to the audio link. For example, the target audio delay includes acquisition delay, preprocessing delay, encoding delay, network transmission delay, decoding delay and playback delay.
[0066] It is understandable that the embodiment of the present application is illustratively described using the transmission of an audio signal as an example, and the transmission of multiple audio signals also applies to a similar principle. That is, if the second device is communicating with multiple devices in real time, the second device can receive audio data packets transmitted by multiple devices. For each audio data packet, delay statistics are required to obtain second delay information, and the second delay information is integrated with the first delay information in the corresponding audio data packet to obtain the delay of the corresponding audio signal as the target audio delay. For example, Figure 3 The data processing scenario diagram shown in the figure shows that in the chorus scenario, in addition to receiving the audio data packet packetA transmitted by the first device, the second device also receives the audio data packet packetB transmitted by the third device. Then, in addition to performing delay statistics based on the audio data packet packetA to obtain the second delay information corresponding to the audio signal in packetA, and integrating the second delay information of packetA with the first delay information recorded in packetA to obtain the delay of the audio signal on the first device side (i.e., target audio delay A), the second device also needs to perform delay statistics based on the audio data packet packetB to obtain the second delay information corresponding to the audio signal in packetB, and integrate the second delay information with the first delay information in the audio data packet packetB to obtain the delay of the audio signal on the third device side (i.e., target audio delay B). In this way, the delay of each audio signal can be obtained, and based on the delay of these audio signals, the end-to-end delay of the audio link can be monitored.
[0067] The data processing method provided by this application can count the end-to-end delay of audio signals transmitted between any two devices during real-time communication. Specifically, it can count the delays generated by the processing stages experienced by the audio signal on different device sides during the transmission of the audio signal throughout the entire audio link. By integrating the delays generated by the audio signal in all processing stages of the audio link, the final target audio delay is obtained. In this way, a detailed and comprehensive end-to-end delay can be obtained, improving the accuracy of the audio delay. For any two devices in the same real-time communication scenario, the above scheme can be used to obtain the corresponding end-to-end delay. That is, no matter how large the real-time communication scenario is, this scheme can cover all devices in the real-time communication scenario for delay statistics, so it is not limited by the scale of use. In addition, no matter how large the differences in device characteristics are, accurate delays can be obtained in real-time usage scenarios. The usage scenarios are not limited to artificially constructed test scenarios. Based on the unlimited usage scale and usage scenarios, it has high versatility and has a good delay statistics rate for large-scale real-time communication scenarios.
[0068] See Figure 4 , is a flow chart of another data processing method provided by an exemplary embodiment of the present application. The data processing method can be performed by a second device (such as Figure 1b The data processing method may be executed by the terminal device 101b) in the embodiment, and may include the following S401-S404.
[0069] S401 , during real-time communication between a first device and a second device, receiving an audio data packet transmitted by the first device.
[0070] In one embodiment, the transmission processing stages in the audio link include one or more of the following: an acquisition stage, a preprocessing stage, and an encoding stage; the first delay information includes: an acquisition delay generated in the acquisition stage, a preprocessing delay generated in the preprocessing stage, and an encoding delay generated in the encoding stage.
[0071] Specifically, the acquisition delay can also be called the first device delay. Since different devices use different parameter configurations and the corresponding acquisition delays vary greatly, the acquisition delay is important for the end-to-end delay of the entire audio link. During the acquisition phase, the first device can acquire audio signals and perform analog-to-digital conversion (A / D conversion) on the acquired audio signals to obtain audio data. Among them, audio data is a data signal, and audio signal is an analog signal. Through analog-to-digital conversion, the analog signal can be converted into a data signal for transmission. Analog-to-digital conversion can include PCM (Pulse Code Modulation). The audio information obtained after digital conversion by PCM can be used as the original audio format. Its essence is a series of digitally converted values of audio signals. A certain delay will be generated in the analog-to-digital conversion process. The first device can count the time consumed by the analog-to-digital conversion process to obtain the first hardware delay. The delays generated by different devices in this process are also different, which is related to the hardware conditions of the first device.
[0072] After analog-to-digital conversion, the first device can also uniformly store the audio data in the acquisition buffer. In this way, a certain amount of audio data will be stored in the acquisition buffer, and the cache will also generate a delay. Therefore, the first device can count the time it takes to cache the audio data and obtain the acquisition buffer delay. During specific statistics, the size of the acquisition buffer can be obtained, and the acquisition buffer delay can be determined based on the size of the acquisition buffer. This is because the size of the cache is directly related to the speed and efficiency of audio data processing. If the cache is small, the data processing speed may not keep up, thereby increasing audio delay; on the contrary, if the cache is large, more data can be stored, making audio processing smoother and the delay relatively low. After obtaining the acquisition buffer delay and the first hardware delay in the acquisition phase, the acquisition delay can be generated based on the acquisition buffer delay and the first hardware delay. The acquisition delay is composed of the acquisition buffer delay and the first hardware delay, and the sum of the two can be used as the acquisition delay. Among them, the first hardware delay refers to the time it takes for the first device to perform analog-to-digital conversion on the collected audio signal during the acquisition phase; the acquisition buffer delay refers to the time it takes for the first device to cache the audio data obtained by analog-to-digital conversion during the acquisition phase.
[0073] For ease of understanding, the following Figure 5a The processing diagram of audio collection is shown in FIG. Figure 5aAs shown, the processing flow includes sound input and audio acquisition, and audio acquisition includes processing items such as sound card A / D, acquisition cache and application acquisition. The first device acquires the sound in the environment where the first device is located and the audio signal based on the voice input by the first device in the sound input link, and performs analog-to-digital conversion on the audio signal in the sound card A / D, and a delay will be generated in this part (which can be called device hardware link delay). The first device can then store the converted audio data in the acquisition cache area in the acquisition cache link, and this part will also generate a delay, namely the acquisition cache delay. Afterwards, the audio data can be used as audio data to be processed in the application acquisition link, and in the subsequent transmission processing stage of the audio link, the first device will perform pre-processing, encoding and other processing on it.
[0074] After the first device obtains the audio data, it can enter the preprocessing stage. In the preprocessing stage, the first device preprocesses the audio data. The main function of this preprocessing is to be responsible for the audio echo cancellation, noise cancellation and audio gain. Since there is also a cache operation in the preprocessing, and the preprocessing cache is generally around 20 milliseconds (ms), depending on the corresponding algorithm. The first device can count the preprocessing delay to obtain the preprocessing delay. The preprocessing delay refers to the time it takes for the first device to perform a cache operation in the process of preprocessing the audio data in the preprocessing stage; the cache operation can be used to cache data that needs to be preprocessed or preprocessed audio data.
[0075] After the preprocessing is completed, the encoding stage begins. In the encoding stage, the first device can use an encoder to encode the preprocessed audio data. Since a certain cache is also stored in the encoder during the encoding process, which is used as a data reference for the encoder during encoding, the cache here will also have a corresponding delay. Specifically, the encoding delay can be obtained by obtaining the cache size of the encoder, and the cache size is used to indicate the cache capacity of the corresponding cache area. Therefore, the encoding delay includes: in the process of encoding the preprocessed audio data in the encoding stage, the length of time it takes to cache the reference data generated in the encoding process, and the reference data is, for example, a reference audio frame.
[0076] After encoding is completed, the uplink phase begins. In the uplink phase, the first device can package the encoded audio stream and, when packaging the audio data, add the acquired acquisition delay, preprocessing delay, and encoding delay to the audio data packet. In addition, in order to obtain the network transmission delay (also known as network delay or network link delay), a data packaging timestamp (i.e., an uplink timestamp) can be added to the audio data packet. The structure of the resulting audio data packet can be as follows.
[0077] Audio packet header Package Type Timestamp Serial number Encoding format Package size Packet Energy Value Acquisition delay Preprocessing delay Encoding delay Uplink timestamp ……
[0078] The structure of the audio data packet described above includes attribute information of the audio data packet, including the audio packet header, packet type, timestamp, sequence number, encoding format, packet size, and packet energy value. The structure of the audio data packet also includes acquisition delay, preprocessing delay, encoding delay, and uplink timestamp. It should be understood that the other information in the audio data packet is not fully illustrated here, and the above structure does not limit the content included in the audio data packet.
[0079] It can be seen that the delay data counted on the first device side are all delays of each link in the uplink. In order to correspond to the downlink and still be a complete link, this application packages the delay of each link in the uplink into the audio data packet, and also packages the timestamp of the data packaging into the audio data packet. Therefore, after the second device receives the audio data packet, based on the processing of the audio data packet, it can obtain the delay on the first device side and determine the network transmission delay in combination with the timestamp on the second device side.
[0080] S402: Perform delay statistics processing on the audio data packet to obtain second delay information of the audio signal.
[0081] In one embodiment, the specific implementation of S402 includes the following contents (1)-(2).
[0082] (1) According to the audio data packet, delay statistics processing is performed at each receiving and processing stage included in the audio link to obtain the delay of the audio signal generated at each receiving and processing stage.
[0083] The various processing stages of the audio link on the second device constitute the downlink, which are referred to as receive processing stages. The logic for performing delay statistics based on audio packets varies in each receive processing stage. The following details the delay statistics logic for each receive processing stage.
[0084] (1.1) The receiving and processing stage includes the unpacking stage, which is as follows Figure 1a In the downlink phase of the audio link shown, the audio data packet also includes a packing timestamp, which refers to the timestamp when the audio data packet is packed. The packing timestamp, also known as the uplink timestamp, is added when the audio data is packed during the uplink phase of the audio link. The packing timestamp can be used to determine network transmission delay. The second device can perform delay statistics based on the audio data packet during the unpacking phase of the audio link to obtain the delay incurred by the audio signal during the unpacking phase. In a specific implementation, the following steps 1.1 to 1.3 can be included.
[0085] Step 1.1: In the unpacking stage, the audio data packet is unpacked to obtain the packing timestamp in the audio data packet.
[0086] During the unpacking phase, when the second device unpacks the audio data packet, it can specifically parse the audio data packet to obtain the uplink timestamp contained in the audio data packet. It is understandable that because the audio data packet also includes first delay information, the unpacking process can also obtain first delay information, which includes acquisition delay, preprocessing delay, and encoding delay. This first delay information can be obtained by the second device and combined with the second delay information statistically obtained on the second device side to determine the target audio delay.
[0087] Step 1.2: Determine the unpacking timestamp of the audio data packet, and obtain the network transmission delay of the audio signal based on the packing timestamp and the unpacking timestamp of the audio data packet.
[0088] In a specific implementation, the unpacking timestamp of an audio data packet refers to the timestamp of the successful unpacking of the audio data packet, which can also be called a downlink timestamp. The absolute value of the difference between the unpacking timestamp and the packing timestamp of the audio data packet can be determined as the network transmission delay. The network transmission delay refers to the time required for audio data to be transmitted in the network, specifically refers to the time it takes for the audio data to be uploaded to the server through the audio data packet and transmitted from the server to the second device for unpacking. Optionally, in order to improve the accuracy of delay statistics, both the unpacking timestamp and the packing timestamp can be accurate to milliseconds, so that the network transmission delay is also accurate to milliseconds.
[0089] Factors affecting network transmission delay may include one or more of the following: data volume, network quality, and device performance. Specifically, data volume refers to the data volume of audio data packets. Network transmission delay is positively correlated with data volume. The larger the audio data packet, the longer it takes to transmit. Network transmission delay is negatively correlated with network quality. The better the network quality, the smaller the network transmission delay, and the worse the network quality, the larger the network transmission delay. Network quality includes but is not limited to: network speed, bandwidth, etc. For example, the higher the network speed, the better the network quality, and the more data can be transmitted per unit time, thus reducing the network transmission delay. Network transmission delay is negatively correlated with device performance. The better the device performance, the smaller the network transmission delay. Device performance here refers to the device used to transmit audio data packets, such as servers, routers or switches in the network, etc. Performance includes but is not limited to: load capacity, storage capacity, and concurrency capacity, etc.
[0090] Step 1.3: The network transmission delay is determined as the delay generated by the audio signal during the depacketization phase.
[0091] Through the above processing, the delay incurred during the unpacking phase of the audio signal now includes network transmission delay. The delay statistics logic during the unpacking phase, described in steps 1.1 through 1.3, accurately and quickly determines the network transmission delay of the audio signal based on the packetization timestamp added by the first device and the unpacking timestamp determined by the second device.
[0092] (2.1) The receiving and processing phase includes a decoding phase. The second device performs delay statistics on the audio data packet during the decoding phase of the audio link to obtain the delay of the audio signal during the decoding phase. In a specific implementation, steps 2.1 through 2.3 may be included.
[0093] Step 2.1: Unpack the audio data packet to obtain the audio stream corresponding to the audio data packet.
[0094] In a specific implementation, the audio data packet also includes the audio signal from the first device, and the audio signal exists in the audio data packet in the form of audio data. The audio data packet includes data of multiple sampling points of the audio data, which constitute an audio stream. Therefore, by unpacking, the data in the audio data packet can be extracted according to a specific format and restored to the audio stream.
[0095] Step 2.2: In the decoding stage, delay statistics processing is performed on the audio stream obtained by the unpacking process to obtain the decoding delay of the audio signal.
[0096] Step 2.3: Determine the decoding delay as the delay generated by the audio signal during the decoding stage.
[0097] In one implementation, when the second device executes step 2.2 above, it may execute the contents described in the following contents ①-③.
[0098] ① During the decoding process of the audio stream, the time spent on caching the decoding reference data generated during the decoding process is counted, and after the decoding process is completed, the time spent on caching the decoding reference data generated during the decoding process is determined as the decoding cache delay of the audio signal.
[0099] Since there is a certain amount of cache in the decoding process, which is used as a reference for decoding, that is, during the decoding process, decoding reference data will be generated and cached. The decoding reference data is, for example, a reference audio frame. The second device can count the time it takes to cache the decoding reference data. The specific statistical methods may include the following two: In one implementation, the size of the corresponding cache area can be obtained for statistics, specifically the cache size and cache speed of the cache area can be obtained; then, based on the ratio between the cache speed and the cache size, the cache delay of the audio signal is obtained. In another implementation, real-time statistics can be performed based on the start / end timestamps. Specifically, the start timestamp when the data is written to the cache area can be determined, and the end timestamp when all the decoding reference data of the audio signal is written to the cache area can be determined. The start timestamp and the end timestamp are subtracted to obtain the time spent on caching the decoding reference data. In this way, real-time statistics based on the start timestamp and the end timestamp can be obtained when the cache is completed.
[0100] ② The audio data obtained by the decoding process is cached. During the cache process, the time spent on the cache process is counted, and the data cache delay of the audio signal is determined based on the time spent on the cache process.
[0101] Specifically, after being unpacked, the downlink audio data packet can first be stored in a downlink cache, such as a data buffer in the jitter buffer, namely the packet buffer (audio packet buffer). The jitter buffer is used to store remote audio frames and provides acceleration and deceleration capabilities to smooth the uneven audio data caused by network jitter, and plays an important role in anti-jitter and weak network conditions. The packet buffer can specifically be used to store received audio data packets. When playback is required, the audio data packets can be obtained from the data buffer (i.e., the packet buffer) for decoding, and acceleration and deceleration operations are performed in sequence. The decoded audio data is then cached in the sync buffer (synchronization buffer). The cache delay is generally between tens of milliseconds and hundreds of milliseconds, so the corresponding cache delay is also not negligible. The second device can obtain the time spent caching the audio data packet, for example, by calculating the corresponding cache delay by obtaining the cache size of the packet buffer. During the process of caching audio data, the time spent caching the audio data can also be counted, for example, by calculating the corresponding cache delay by obtaining the cache size of the sync buffer. The time spent caching audio data can be used as the cache delay of the audio data. The data cache delay can be obtained by summing the cache delay of the audio data and the cache delay of the audio data packet. That is, the data cache delay is composed of the cache delay of the audio data and the cache delay of the audio data packet.
[0102] ③ Integrate the data buffer delay of the audio signal and the decoding buffer delay of the audio signal to obtain the decoding delay of the audio signal.
[0103] In a specific implementation, the second device can sum the above-stated data cache delay and decoding cache delay to obtain the decoding delay, so that the decoding delay is composed of the data cache delay and the decoding cache delay. It should be noted that the time consumed by the various processes involved in the decoding processing stage (such as transformation, quantization, etc.) can be ignored compared to the cache delay due to the fast processing speed. Therefore, the statistical logic of the decoding delay shown in ①-③ above is mainly to count the time spent on the cache of various data involved in the decoding stage, so that the decoding delay obtained is composed of multiple cache delays. In order to perform statistics more comprehensively, the time spent on each process involved in the decoding process can also be counted to obtain the decoding processing delay, so that the decoding delay is composed of the cache delay and the decoding processing delay.
[0104] (3.1) The receiving and processing phase includes a playback phase. The second device can perform delay statistics based on the audio data packet during the playback phase of the audio link to obtain the delay of the audio signal during the playback phase. In a specific implementation, the following steps 3.1 to 3.3 can be included.
[0105] Step 3.1: Decode the audio data packet to obtain audio data corresponding to the audio data packet.
[0106] Specifically, the audio data packet may be first unpacked to obtain an audio stream, and then the audio stream may be decoded to obtain audio data. It is understood that during both the unpacking and decoding processes, relevant delay data is collected to obtain the delay of the audio signal at the corresponding processing stage.
[0107] Step 3.2: During the playback phase, perform delay statistics processing based on the audio data to obtain the playback delay of the audio signal.
[0108] During the playback phase, the audio data will be converted into an audio signal, that is, the digital signal will be converted into an analog signal. These audio signals will be cached in the playback buffer area as playback data, and then played on the second device. Therefore, the playback delay obtained by delay statistics is composed of the playback buffer delay and the second hardware delay. The playback buffer delay refers to the time it takes to cache the audio data, and the second hardware delay refers to the time it takes for the second device to convert the audio data based on its digital-to-analog conversion capability. Since these two delays are related to the performance of the second device itself, the playback delay can also be called the second device delay.
[0109] In one embodiment, the number of audio data packets includes one, and the decoding process obtains one audio data packet. When the second device executes step 3.2, the steps may specifically include the following contents shown in ①-③.
[0110] ① During the playback phase, the audio data is subjected to digital-to-analog conversion processing, and during the digital-to-analog conversion process, the duration of the digital-to-analog conversion process is counted to obtain the second hardware delay of the audio signal.
[0111] In a specific implementation, when in the playback phase, it indicates that the audio data packet has been unpacked and decoded to obtain audio data. In order to play the audio signal in the second device, the second device can first perform digital-to-analog conversion (i.e., Digital to Analog, D / A conversion) on the audio data to obtain the audio signal. The digital-to-analog conversion process is the inverse process of the analog-to-digital conversion process, and its function is to convert the digital signal (corresponding to the audio data) into an analog signal (corresponding to the audio signal). During the conversion process, since the digital-to-analog conversion is related to the D / A of the sound card of the second device itself, this process has a certain delay, and there are also differences between different devices. Therefore, during the digital-to-analog conversion process, the second device can count the time spent on the digital-to-analog conversion of the audio data, and after the conversion process is completed, the time spent on the digital-to-analog conversion process is determined as the second hardware delay. Specifically, the start timestamp and the end timestamp of the digital-to-analog conversion of the audio data can be subtracted to obtain the second hardware delay. In one implementation, since there is also a cache operation in the digital-to-analog conversion process, the cache operation can cache the audio data that needs to be digital-to-analog converted. The second device can also count the time consumed by the cache operation and add the time consumed to the second hardware delay, so that the second hardware delay includes the time consumed by the cache operation.
[0112] ② Performing a cache process on the audio signal. During the cache process, the duration of caching the audio signal is counted, and the duration of caching the audio signal is determined as the playback cache delay of the audio signal.
[0113] The audio signal obtained after the audio data conversion is completed can be uniformly stored in the playback cache, and the size of the playback cache will also affect the playback of the audio signal in the second device. Therefore, during the cache processing, it is also necessary to count the time spent on the cache, so that the time spent on caching the audio signal can be determined as the playback cache delay. For the statistical method of the playback cache delay, you can also refer to the other cache delay related content introduced above, for example, first obtain the cache size of the playback cache area, and then determine the playback cache delay based on the cache size.
[0114] ③ Integrate the playback buffer delay of the audio signal and the second hardware delay of the audio signal to obtain the playback delay of the audio signal.
[0115] Specifically, the playback buffer delay and the second hardware delay can be summed to obtain the playback delay of the audio signal. That is, the playback delay of the audio signal is composed of the playback buffer delay and the second hardware delay. Since the playback buffer and digital-to-analog conversion are both related to the second device itself, the playback delay can also be understood as the delay caused by the second device in the process of processing audio data (referred to as the second device delay). For different devices, the playback delay corresponding to different parameter configurations has certain differences.
[0116] Step 3.3: Determine the playback delay as the delay generated by the audio signal during the playback phase.
[0117] Specifically, the delay generated by the audio signal during the playback phase is the playback delay. Figure 5b The process flow diagram of device playback shown in the figure introduces how to obtain playback delay. Figure 5b As shown, the processing flow mainly includes: sound output and audio playback. The audio playback includes the sound card D / A, playback buffer, and application playback. Among them, the sound card D / A and playback buffer are the parts with delay. Sound output can obtain audio data, and then send the audio data to the sound card D / A. After conversion processing by the sound card D / A, it can be converted into an analog signal and stored in the playback buffer area through the playback buffer. In the processing flow of the sound card D / A and playback buffer, accurate delay can be obtained based on the processing of the audio signal. Afterwards, the audio signal cached in the playback buffer can be played on the second device through the application playback.
[0118] In the contents described in steps 3.1 to 3.3 above, the delay can be counted in real time during the processing of the audio data. Based on the processing performed by the second device for playing the audio signal, the impact of the second device on the delay of the entire link can be determined in detail.
[0119] In one embodiment, the number of audio data packets includes multiple audio data packets, and the multiple audio data packets include audio data packets transmitted by the first device and audio data packets transmitted by other devices, and real-time communication is performed between the other devices and the first device. Therefore, each received audio data packet is unpacked and decoded, and multiple audio data are obtained through decoding, and one audio data corresponds to one audio data packet. That is, the second device can receive audio data packets transmitted by different devices, so that the second device can obtain multiple downstream audio streams. The receiving and processing stage also includes: a mixing stage, which is before the playing stage, that is, multiple audio data need to be mixed, and then the mixed data obtained by mixing is played through the playing stage. Based on this, before executing step 3.2, the following steps 4.1-step 4.3 can also be executed.
[0120] Step 4.1: In the stream mixing stage, multiple audio data are mixed to obtain mixed stream data.
[0121] Specifically, when multiple audio streams appear, a stream mixing operation needs to be performed. The stream mixing operation can be used to superimpose the various audio streams, so that the audio stream generated contains information about multiple audio streams. Therefore, the mixed stream data obtained by the second device by mixing the audio data includes information about multiple audio data. The mixed stream data is used to determine the playback delay of the audio signal during the playback phase. Specifically, during the playback phase, the mixed stream data can be digital-to-analog converted, and during the digital-to-analog conversion process, the second device delay is statistically obtained. As a result, the second device delay is the same for each audio signal. After the mixed stream data is processed by digital-to-analog conversion, a mixed stream signal can be obtained. The mixed stream signal is stored in the playback buffer area, so that the delay caused by the mixed stream signal cache can be obtained as the playback buffer delay of each audio signal. Finally, the playback delay is obtained based on the second device delay and the playback buffer delay. After multiple audio data are mixed, the mixing delay and playback delay of each audio signal are the same.
[0122] Step 4.2: During the stream mixing process, the time spent on the stream mixing process is counted, and after the stream mixing process is completed, the time spent on the stream mixing process is determined as the stream mixing delay of each audio signal.
[0123] In the process of mixing the audio data, the time from the start to the end of the mixing process can be calculated and determined as the mixing delay. Since the mixing process is to superimpose the audio data, the mixing delay is the same for each audio signal.
[0124] Step 4.3: Determine the mixing delay of each audio signal as the delay generated by the corresponding audio signal during the mixing stage.
[0125] For audio signals on different device sides, the same mixing delay can be obtained in the mixing stage, and the playback delay generated in the subsequent playback stage can be unified. However, since the delay generated on the first device side indicated by the first delay information may be different, for multi-channel audio signals, even if the playback delay and the mixing delay are the same, the state of synchronous playback of multi-channel audio signals may not be achieved. When multiple audio streams are simultaneously downlinked, by processing the audio delay of multiple audio streams, the above process can be applied to the scenario of multiple audio streams and solve the delay problem of multiple audio streams. In this way, the delay difference of different audio streams is reduced, and multiple audio streams are played synchronously as much as possible.
[0126] (2) The delays generated by the audio signal in each receiving and processing stage are integrated to obtain the second delay information of the audio signal.
[0127] In one implementation, the second device may combine the delays generated by the audio signal in each receiving and processing stage to obtain a first delay set, which includes the delays generated by the audio signal in each receiving and processing stage; and then determine the first delay set as the second delay information of the audio signal. In another implementation, the second device may sum the delays generated by the audio signal in each receiving and processing stage to obtain a total delay value, and then use the total delay value as the second delay information of the audio signal. For example, based on the delays generated in the unpacking stage, decoding stage, and playback stage introduced above, if the second device only receives one audio data packet, then for the audio signal corresponding to the audio data packet, the calculation formula for the downlink delay can be as follows: downlink delay = playback delay + network transmission delay + decoding delay. If the second device receives m (m is an integer) audio data packets, and these m audio data packets are transmitted by different devices, then for the i-th (i∈[1,m]) audio signal, the downlink delay is calculated as follows: Downlink delay i = Playback delay i + Network transmission delay i + Decoding delay i + Mixing delay i, where i represents the i-th downlink audio stream.
[0128] The contents described in (1)-(2) above can obtain the delay of each stage in the downlink by statistically analyzing the delay generated by the audio signal in each receiving and processing stage, and can cover the scenario of multiple audio streams, so as to obtain the downlink delay more comprehensively and accurately.
[0129] Based on the description of the above content, this application Figure 1a The audio link shown is based on the provided Figure 6 The diagram of audio link delay is shown in Figure 1. Figure 6As shown, in the audio link, there are corresponding delays in each stage, including: the acquisition delay in the audio acquisition stage, the 3A delay in the 3A stage (i.e., preprocessing delay), and the encoding delay in the encoding stage. When the data is packaged, an uplink timestamp will be added, and the audio data packet will be uploaded to the server. When the device on the near-end side downloads the audio data packet, the downlink timestamp can be determined when the data is unpacked, so that the network transmission delay can be obtained based on the uplink timestamp and the downlink timestamp. Therefore, it also includes: uplink and downlink network transmission delays, decoding delays in the decoding stage, mixing delays in the mixing stage, and playback delays in the playback stage. Among them, the mixing stage can calculate multiple delays to obtain the mixing delay. It can be seen that the device on the near-end side can obtain the delays of each stage in the complete audio link and can process the audio delays of multiple downlinks. It can be understood that for the mixing stage, if only one downlink data is obtained, the mixing stage can be disabled, that is, the decoded audio data is still output as audio data after the mixing stage. Optionally, the corresponding device can identify the links in the audio link that may have delays, specifically the links that involve caching. That is, if data is cached in a link, it indicates that this link will have delays, so delay statistics can be collected for the links with delays. As shown above, in the audio link, almost every link will have delays.
[0130] In one embodiment, the first device and the second device communicate in real time based on the target application, and the audio data packets transmitted by the real-time communication between the first device and the second device are forwarded by the application server of the target application, that is, the audio data packets packaged by the first device can be uploaded to the application server, and then downloaded to the second device by the application server. The application server can obtain the delay of the audio signal on the first device side (i.e., the first delay information) based on the analysis of the audio data packet. Therefore, after obtaining the second delay information of the audio signal, the second device can also: send the second delay information of the audio signal to the application server of the target application, so that the application server can check the delay of the audio link based on the second delay information and the first delay information.
[0131] In one implementation, each processing stage in the audio link is used to indicate the processing logic for the audio signal. The processing logic is used to describe the processing content performed on the audio signal or the parameters, algorithms, etc. used. For example, the processing stage is the encoding stage, which indicates the algorithm used to encode the audio signal, and the processing stage is the preprocessing stage, which indicates the parameter settings for preprocessing the audio signal. Since the processing stages in the audio link are divided into the transmission processing stage and the reception processing stage, the type of any processing stage is the transmission processing stage or the reception processing stage. Delay troubleshooting includes at least one of the following ①-③:
[0132] ① If the delay generated by the audio signal in any processing stage is greater than or equal to the first preset delay, the processing logic indicated by the corresponding processing stage is optimized. Specifically, the first preset delay is a delay used to ensure that the delay generated by the audio signal processing process is within an acceptable range. The first preset delay can be set based on experience, or based on the delay generated by the processing process of historical audio signals. Assuming that the audio link includes P (P is an integer) processing stages, if the delay generated by the audio signal in processing stage p (p∈[1, P], and p is an integer) is greater than or equal to the first preset delay, it means that the delay generated by the audio signal in processing stage p may cause a large end-to-end delay, so that the application server can optimize the processing logic indicated by processing stage p to control the delay generated by the audio signal in processing stage p within the first preset delay. During optimization, the application server can perform optimization that matches the processing logic indicated by processing stage p. For example, if the processing stage p is the encoding stage, then one may consider adjusting to a better-performing encoding standard, or adopting a more optimized audio processing algorithm, or optimizing the encoding buffer size, etc. In another implementation, a first preset delay matching the processing stage may be set for each processing stage, so that the delay generated by each processing stage can be compared with the matching first preset delay to more accurately determine whether the processing stage needs to be optimized.
[0133] ② If the sum of the delays generated by the audio signal in each transmission processing stage is greater than or equal to the second preset delay, the transmission processing stage with the largest delay is selected from each transmission processing stage as the target processing stage, and the processing logic indicated by the target processing stage is optimized.
[0134] Specifically, the application server can sum the delays generated by the audio signal indicated by the first delay information in each transmission processing stage to obtain a sum value as the uplink delay. The uplink delay is then compared with the second preset delay. The second preset delay is a value that ensures that the uplink delay is within an acceptable range, and its setting method can be similar to the setting of the first preset delay. If the uplink delay is greater than or equal to the second preset delay, it means that the accumulated delays of each link in the uplink are large, which will result in a larger end-to-end delay. Therefore, in order to reduce the delay, the transmission processing stage with the largest delay can be selected from each transmission stage as the target processing stage, and the processing logic of the target processing stage can be optimized to give priority to reducing the delay generated by the target processing stage, so that the uplink delay is controlled within the second preset delay, thereby reducing the end-to-end delay as much as possible.
[0135] ③ If the sum of the delays generated by the audio signal in each receiving and processing stage is greater than or equal to the third preset delay, the receiving and processing stage with the largest delay is selected from each receiving and processing stage as the target processing stage, and the processing logic indicated by the target processing stage is optimized.
[0136] Specifically, the application server can sum the delays generated by the audio signal indicated by the first delay information in each transmission processing stage to obtain a sum value as the downlink delay, which includes the network transmission delay. Similar to the transmission processing stage, when the downlink delay is greater than or equal to the third preset delay, the application server can select the receiving processing stage with the largest delay. For example, the delays of each receiving processing stage can be sorted from large to small, and then the receiving processing stage with the highest delay sorting can be used as the target processing stage. The application server can optimize the processing logic of the target processing stage to prioritize reducing the delay generated by the target processing stage, so that the downlink delay is controlled within the third preset delay, thereby reducing the end-to-end delay as much as possible.
[0137] The above-mentioned link delay troubleshooting can identify the processing stage that needs to be optimized in the audio link by comparing the size between a single processing stage and the first preset delay, or by comparing the size between the sum of the delays of multiple processing stages and the corresponding preset delay. By optimizing the processing logic indicated by the corresponding processing stage, the delay of the processing stage can be reduced, thereby reducing the end-to-end delay of the audio link.
[0138] S403: Integrate the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay.
[0139] When receiving multiple channels of downlink data (i.e., audio data packets sent by different devices), the delay of each audio stream before mixing can be different, and the multi-channel delay calculation link can calculate the delay of each audio stream separately. In one implementation, the first delay information includes acquisition delay, preprocessing delay, and encoding delay, and the second delay information includes network transmission delay, decoding delay, mixing delay, and playback delay. Then, the delay of each stage of the corresponding audio stream in the uplink and the delay of the downlink stage (including network transmission delay) can be accumulated to obtain the delay of each audio stream before mixing. That is: delay before mixing i = acquisition delay i + preprocessing delay i + encoding delay i + network transmission delay i + decoding delay i (including buffer delay), where i represents the i-th downlink audio stream (i.e., the i-th audio signal). For different audio signals, due to differences between different devices, the delay before mixing of different audio signals can be different. On the second device, since all audio signals share the same playback delay and mixing delay after mixing, the delay calculation for the i-th audio signal across the entire audio link is as follows: Target audio delay = playback delay + mixing delay + pre-mixing delay i. When receiving a downlink data stream (i.e., the audio data packet sent by the first device), since mixing is not required, the target audio delay = playback delay + acquisition delay + pre-processing delay + encoding delay + network transmission delay + decoding delay.
[0140] It can be seen that the application can obtain detailed information on the delay of each audio stream corresponding to the downlink, and can also obtain the delay of the entire audio link. In this way, the overall end-to-end delay situation can be presented, and it can also be convenient for application manufacturers to improve and monitor the delay of the target application to reduce the delay as much as possible and improve real-time performance. In one implementation method, the second device can also send the target audio delay to the application server of the target application, so that the application server can also perform targeted optimization of the target application based on the target audio delay. In this way, backup data can be provided for the basis of targeted optimization to avoid the inability to obtain accurate end-to-end delay and complete delay data when the first delay information of the corresponding audio signal is lost.
[0141] In one embodiment, the first device and the second device perform real-time communication based on a target application, and the target audio delay is obtained by summing the first delay information of the audio signal and the second delay information of the audio signal. After obtaining the target audio delay, the following step S404 may be performed.
[0142] S404: If the target audio delay is greater than or equal to the preset audio delay, or the number of audio signals whose target audio delay is greater than or equal to the preset audio delay exceeds a preset threshold, a delay prompt is output.
[0143] In the case where the target audio delay is the end-to-end delay of the audio link, if the second device receives an audio data packet, the target audio delay can be compared with the preset audio delay. If the target audio delay is greater than the preset audio delay, it indicates that the end-to-end delay is large, which will affect the real-time performance of the audio signal playback. Therefore, the second device can output a delay prompt to prompt the user to check the operating settings of the target application and adjust it if necessary. If the second device receives multiple audio data packets sent by different devices, the audio signal corresponding to each audio data packet can be used to determine the target audio delay. Therefore, the target audio delay of each audio signal can be compared with the preset audio delay, and the number of audio signals greater than or equal to the preset audio delay can be counted. If the number obtained is greater than or equal to the preset number threshold, it indicates that the operating settings of the target application in the second device may be the cause of the large end-to-end delay, and thus a delay prompt can be output. In this way, by comparing the target audio delays of multiple audio signals with the preset audio delays, it is possible to accurately determine whether to output a delay prompt based on the comparison results of the multiple audio signals.
[0144] In one feasible manner, the delayed prompt is used to prompt the adjustment of the running settings of the target application. The delayed prompt can be one or more of text, image, video or voice, and this application does not impose any restrictions on this. The target application is integrated with audio acquisition function and data transmission function. The audio acquisition function can obtain the user's voice input on the first device side, the background sound on the first device side, etc., and after the audio signal is acquired, it needs to be transmitted to other devices through the network. Considering the transmission speed and network bandwidth, the audio signal is encoded and compressed on the first device side before being transmitted. Based on this, the running settings of the target application include at least one of the following: the frequency of collecting audio signals; the rate of transmitting audio signals. In addition, since the audio signal needs to be pre-processed (such as echo cancellation, noise cancellation, etc.), encoded, etc., the running settings may also include one or more of the following: audio coding standards, sound gain size, etc. Among them, the audio coding standards may include: AVS related standards, such as AVS2, etc.
[0145] The data processing method provided in the embodiment of the present application can be used to count the end-to-end delay of the audio signal transmitted between the two devices during the real-time communication between any two devices. Specifically, it can be used to count the delays generated by the processing stages experienced by the audio signal on different device sides during the transmission of the audio signal throughout the entire audio link. By integrating the delays generated by the audio signal in all processing stages of the audio link, the final target audio delay is obtained. In this way, a detailed and comprehensive end-to-end delay can be obtained, and the accuracy of the audio delay can be improved. Moreover, for any device (second device) that receives an audio signal, it can receive audio signals transmitted by multiple other devices simultaneously or successively, and mix the audio signals, and count the delays in the mixing stage. This can solve the problem of delay of multiple audio streams and further improve the accuracy and comprehensiveness of the end-to-end delay of the audio link. Furthermore, after obtaining the target audio delay, the audio link can be optimized based on the target audio delay or the settings of the target application used to support real-time communication, so as to reduce the end-to-end delay and improve the real-time performance in the real-time communication scenario, thereby improving the user experience.
[0146] See Figure 7 , Figure 7 This is a structural diagram of a data processing device provided in an embodiment of the present application. The data processing device can be set in the computer device provided in an embodiment of the present application. Figure 7 The data processing device shown may be a computer program (including program code) running in a computer device, which may be used to perform Figure 2 or Figure 4 Some or all of the steps in the method embodiment shown. Figure 7 , the data processing device may include the following units:
[0147] A receiving unit 701 is configured to receive an audio data packet transmitted by the first device during real-time communication between the first device and the second device; the audio data packet includes an audio signal from the first device and first delay information, the first delay information being used to indicate delays incurred at various transmission processing stages on the first device during transmission of the audio signal along the audio link.
[0148] The processing unit 702 is configured to perform delay statistical processing on the audio data packet to obtain second delay information of the audio signal; the second delay information is used to indicate the delays incurred by the audio signal at various receiving and processing stages experienced by the second device during transmission along the audio link.
[0149] The processing unit 702 is further configured to integrate the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay.
[0150] In one embodiment, the processing unit 702 is specifically used to: perform delay statistical processing at each receiving and processing stage included in the audio link according to the audio data packet to obtain the delay generated by the audio signal at each receiving and processing stage; and integrate the delay generated by the audio signal at each receiving and processing stage to obtain second delay information of the audio signal.
[0151] In one embodiment, the receiving and processing stage includes an unpacking stage, and the audio data packet also includes a packing timestamp. The processing unit 702 is specifically used to: unpack the audio data packet in the unpacking stage to obtain the packing timestamp in the audio data packet; determine the unpacking timestamp of the audio data packet, and obtain the network transmission delay of the audio signal based on the packing timestamp and the unpacking timestamp of the audio data packet; and determine the network transmission delay as the delay generated by the audio signal in the unpacking stage.
[0152] In one embodiment, the receiving and processing stage includes a decoding stage. Based on the audio data packet, the processing unit 702 is specifically used to: unpack the audio data packet to obtain an audio stream corresponding to the audio data packet; in the decoding stage, perform delay statistics processing on the audio stream obtained by the unpacking process to obtain a decoding delay of the audio signal; and determine the decoding delay as the delay generated by the audio signal in the decoding stage.
[0153] In one embodiment, in the decoding stage, the processing unit 702 is specifically used to: during the decoding process of decoding the audio stream in the decoding stage, count the time spent on caching the decoding reference data generated in the process of decoding, and after the decoding process is completed, determine the time spent on caching the decoding reference data generated in the process of decoding as the decoding cache delay of the audio signal; cache the audio data obtained by the decoding process, and during the cache process, count the time spent on the cache process, and determine the data cache delay of the audio signal based on the time spent on the cache process; integrate the data cache delay of the audio signal and the decoding cache delay of the audio signal to obtain the decoding delay of the audio signal.
[0154] In one embodiment, the receiving and processing stage includes a playback stage, and the processing unit 702 is specifically used to: perform decoding processing based on the audio data packet to obtain audio data corresponding to the audio data packet; in the playback stage, perform delay statistical processing based on the audio data to obtain the playback delay of the audio signal; and determine the playback delay as the delay generated by the audio signal during the playback stage.
[0155] In one embodiment, the number of audio data packets includes one, and the decoding process obtains one audio data; the processing unit 702 is specifically used to: in the playback stage, perform digital-to-analog conversion on the audio data, and during the digital-to-analog conversion process, count the time spent on the digital-to-analog conversion process to obtain a second hardware delay of the audio signal; perform cache processing on the audio signal, and during the cache processing process, count the time spent on caching the audio signal, and determine the time spent on caching the audio signal as the playback cache delay of the audio signal; integrate the playback cache delay of the audio signal and the second hardware delay of the audio signal to obtain the playback delay of the audio signal.
[0156] In one embodiment, the number of audio data packets includes multiple audio data packets, and the multiple audio data packets include audio data packets transmitted by the first device and audio data packets transmitted by other devices; the decoding process obtains multiple audio data, and the receiving processing stage includes a mixing stage; in the playback stage, delay statistics processing is performed according to the audio data to obtain the playback delay of the audio signal, and the processing unit 702 is also used to: in the mixing stage, mix the multiple audio data to obtain mixed data; the mixed data is used to determine the playback delay of the audio signal in the playback stage; in the process of mixing, the time spent on the mixing process is counted, and after the mixing process is completed, the time spent on the mixing process is determined as the mixing delay of each audio signal; the mixing delay of each audio signal is determined as the delay generated by the corresponding audio signal in the mixing stage.
[0157] In one embodiment, the processing unit 702 is specifically configured to combine delays generated by the audio signal in each receiving and processing stage to obtain a first delay set, and determine the first delay set as the second delay information of the audio signal.
[0158] In one embodiment, the transmission processing stage includes one or more of the following: an acquisition stage, a preprocessing stage, and an encoding stage, and the first delay information includes: an acquisition delay generated in the acquisition stage, a preprocessing delay generated in the preprocessing stage, and an encoding delay generated in the encoding stage; the acquisition delay is composed of an acquisition buffer delay and a first hardware delay, wherein the first hardware delay refers to the time it takes for the first device to perform analog-to-digital conversion on the acquired audio signal in the acquisition stage; the acquisition buffer delay refers to the time it takes for the first device to cache the audio data obtained by the analog-to-digital conversion in the acquisition stage; the preprocessing delay refers to the time it takes for the first device to perform a cache operation on the audio data in the preprocessing stage; the encoding delay includes: the time it takes to cache the reference data generated in the encoding process during the encoding process of the preprocessed audio data in the encoding stage.
[0159] In one embodiment, the processing unit 702 is specifically configured to: combine the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay; or sum the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay.
[0160] In one embodiment, the first device and the second device communicate in real time through the target application, and the audio data packet is sent to the second device through the application server of the target application; after delay statistics processing is performed on the audio data packet to obtain second delay information of the audio signal, the sending unit 703 is used to: send the second delay information of the audio signal to the application server of the target application, so that the application server can perform delay check on the audio link based on the second delay information and the first delay information.
[0161] In one embodiment, each processing stage is used to indicate the corresponding processing logic for the audio signal, and the type of any processing stage is a transmission processing stage or a receiving processing stage; the delay troubleshooting includes at least one of the following: if the delay generated by the audio signal in any processing stage is greater than or equal to a first preset delay, the processing logic indicated by the corresponding processing stage is optimized; or, if the sum of the delays generated by the audio signal in each transmission processing stage is greater than or equal to a second preset delay, the transmission processing stage with the largest delay is selected from each transmission processing stage as the target processing stage, and the processing logic indicated by the target processing stage is optimized; or, if the sum of the delays generated by the audio signal in each receiving processing stage is greater than or equal to a third preset delay, the receiving processing stage with the largest delay is selected from each receiving processing stage as the target processing stage, and the processing logic indicated by the target processing stage is optimized.
[0162] In one embodiment, the first device and the second device communicate in real time through a target application, and the target audio delay is obtained by summing the first delay information of the audio signal and the second delay information of the audio signal. The output unit 704 is used to: if the target audio delay is greater than or equal to the preset audio delay, or the number of audio signals whose target audio delay is greater than or equal to the preset audio delay exceeds a preset number threshold, then output a delay prompt; wherein the delay prompt is used to prompt adjustment of the operating settings of the target application; the operating settings include at least one of the following: the frequency of collecting audio signals; the rate of transmitting audio signals.
[0163] In one embodiment, the real-time communication between the first device and the second device includes any one of the following: real-time communication in a live broadcast scenario, real-time communication in an audio call scenario, real-time communication in a video call scenario, real-time communication in a chorus scenario, real-time communication in a game scenario, and real-time communication in an online meeting scenario.
[0164] It is understood that the specific functions of the various units of the data processing device described in the embodiments of the present application can be specifically implemented according to the methods in the above-mentioned method embodiments. The specific implementation process can refer to the relevant description of the above-mentioned method embodiments and will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated.
[0165] The present application also provides a schematic diagram of the structure of a computer device. The schematic diagram of the structure of the computer device can be found in Figure 8 The computer device may include a processor 801, an input device 802, an output device 803, and a memory 804. The processor 801, input device 802, output device 803, and memory 804 are connected via a bus. The memory 804 is configured to store a computer-readable storage medium including a computer program. The processor 801 is configured to execute the computer program stored in the memory 804.
[0166] In one embodiment, the processor 801 performs the following operations by running a computer program in the memory 804: during real-time communication between the first device and the second device, receiving an audio data packet transmitted by the first device; the audio data packet includes an audio signal and first delay information on the first device side, and the first delay information is used to indicate: the delay caused by the audio signal in each transmission processing stage experienced by the first device side during the transmission according to the audio link; performing delay statistical processing on the audio data packet to obtain second delay information of the audio signal; the second delay information is used to indicate: the delay caused by the audio signal in each receiving processing stage experienced by the second device side during the transmission according to the audio link; integrating the first delay information of the audio signal and the second delay information of the audio signal to obtain the target audio delay.
[0167] In one embodiment, the processor 801 is specifically used to: perform delay statistical processing in each receiving and processing stage included in the audio link according to the audio data packet to obtain the delay generated by the audio signal in each receiving and processing stage; integrate the delay generated by the audio signal in each receiving and processing stage to obtain second delay information of the audio signal.
[0168] In one embodiment, the receiving and processing stage includes an unpacking stage, and the audio data packet also includes a packing timestamp. The processor 801 is specifically used to: in the unpacking stage, unpack the audio data packet to obtain the packing timestamp in the audio data packet; determine the unpacking timestamp of the audio data packet, and obtain the network transmission delay of the audio signal based on the packing timestamp and the unpacking timestamp of the audio data packet; determine the network transmission delay as the delay generated by the audio signal in the unpacking stage.
[0169] In one embodiment, the receiving and processing stage includes a decoding stage. Based on the audio data packet, the processor 801 is specifically used to: unpack the audio data packet to obtain an audio stream corresponding to the audio data packet; in the decoding stage, perform delay statistics processing on the audio stream obtained by the unpacking process to obtain a decoding delay of the audio signal; and determine the decoding delay as the delay generated by the audio signal in the decoding stage.
[0170] In one embodiment, in the decoding stage, the processor 801 is specifically used to: during the decoding process of decoding the audio stream in the decoding stage, count the time spent on caching the decoding reference data generated in the process of decoding, and after the decoding process is completed, determine the time spent on caching the decoding reference data generated in the process of decoding as the decoding cache delay of the audio signal; cache the audio data obtained by the decoding process, and during the cache process, count the time spent on the cache process, and determine the data cache delay of the audio signal based on the time spent on the cache process; integrate the data cache delay of the audio signal and the decoding cache delay of the audio signal to obtain the decoding delay of the audio signal.
[0171] In one embodiment, the receiving and processing stage includes a playback stage, and the processor 801 is specifically used to: decode the audio data packet to obtain audio data corresponding to the audio data packet; in the playback stage, perform delay statistical processing on the audio data to obtain the playback delay of the audio signal; and determine the playback delay as the delay generated by the audio signal during the playback stage.
[0172] In one embodiment, the number of audio data packets includes one, and the decoding process obtains one audio data; the processor 801 is specifically used to: in the playback stage, perform digital-to-analog conversion on the audio data, and during the digital-to-analog conversion process, count the time spent on the digital-to-analog conversion process to obtain a second hardware delay of the audio signal; cache the audio signal, and during the cache process, count the time spent on caching the audio signal, and determine the time spent on caching the audio signal as the playback cache delay of the audio signal; integrate the playback cache delay of the audio signal and the second hardware delay of the audio signal to obtain the playback delay of the audio signal.
[0173] In one embodiment, the number of audio data packets includes multiple audio data packets, and the multiple audio data packets include audio data packets transmitted by the first device and audio data packets transmitted by other devices; the decoding process obtains multiple audio data, and the receiving processing stage includes a mixing stage; in the playback stage, delay statistics processing is performed according to the audio data to obtain the playback delay of the audio signal, and the processor 801 is also used to: in the mixing stage, mix the multiple audio data to obtain mixed data; the mixed data is used to determine the playback delay of the audio signal in the playback stage; in the process of mixing, the time spent on the mixing process is counted, and after the mixing process is completed, the time spent on the mixing process is determined as the mixing delay of each audio signal; the mixing delay of each audio signal is determined as the delay generated by the corresponding audio signal in the mixing stage.
[0174] In one embodiment, the processor 801 is specifically configured to: combine delays generated by the audio signal in each receiving and processing stage to obtain a first delay set, and determine the first delay set as the second delay information of the audio signal.
[0175] In one embodiment, the transmission processing stage includes one or more of the following: an acquisition stage, a preprocessing stage, and an encoding stage, and the first delay information includes: an acquisition delay generated in the acquisition stage, a preprocessing delay generated in the preprocessing stage, and an encoding delay generated in the encoding stage; the acquisition delay is composed of an acquisition buffer delay and a first hardware delay, wherein the first hardware delay refers to the time it takes for the first device to perform analog-to-digital conversion on the acquired audio signal in the acquisition stage; the acquisition buffer delay refers to the time it takes for the first device to cache the audio data obtained by the analog-to-digital conversion in the acquisition stage; the preprocessing delay refers to the time it takes for the first device to perform a cache operation on the audio data in the preprocessing stage; the encoding delay includes: the time it takes to cache the reference data generated in the encoding process during the encoding process of the preprocessed audio data in the encoding stage.
[0176] In one embodiment, the processor 801 is specifically configured to: combine the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay; or sum the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay.
[0177] In one embodiment, a first device and a second device communicate in real time through a target application, and an audio data packet is sent to the second device through an application server of the target application; after performing delay statistical processing on the audio data packet to obtain second delay information of the audio signal, the processor 801 is used to: send the second delay information of the audio signal to the application server of the target application, so that the application server performs delay check on the audio link based on the second delay information and the first delay information.
[0178] In one embodiment, each processing stage is used to indicate the corresponding processing logic for the audio signal, and the type of any processing stage is a transmission processing stage or a receiving processing stage; the delay troubleshooting includes at least one of the following: if the delay generated by the audio signal in any processing stage is greater than or equal to a first preset delay, the processing logic indicated by the corresponding processing stage is optimized; or, if the sum of the delays generated by the audio signal in each transmission processing stage is greater than or equal to a second preset delay, the transmission processing stage with the largest delay is selected from each transmission processing stage as the target processing stage, and the processing logic indicated by the target processing stage is optimized; or, if the sum of the delays generated by the audio signal in each receiving processing stage is greater than or equal to a third preset delay, the receiving processing stage with the largest delay is selected from each receiving processing stage as the target processing stage, and the processing logic indicated by the target processing stage is optimized.
[0179] In one embodiment, a first device and a second device communicate in real time through a target application, and a target audio delay is obtained by summing the first delay information of the audio signal and the second delay information of the audio signal. The processor 801 is used to: output a delay prompt if the target audio delay is greater than or equal to the preset audio delay, or the number of audio signals whose target audio delay is greater than or equal to the preset audio delay exceeds a preset number threshold; wherein the delay prompt is used to prompt adjustment of the operating settings of the target application; the operating settings include at least one of the following: the frequency of collecting audio signals; the rate of transmitting audio signals.
[0180] In one embodiment, the real-time communication between the first device and the second device includes any one of the following: real-time communication in a live broadcast scenario, real-time communication in an audio call scenario, real-time communication in a video call scenario, real-time communication in a chorus scenario, real-time communication in a game scenario, and real-time communication in an online meeting scenario.
[0181] It should be understood that the computer device described in the embodiments of the present application can execute the description of the data processing method in the corresponding embodiments above, and can also execute the description of the data processing device in the corresponding embodiments above, which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here.
[0182] In addition, it should be noted that the present invention also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the above Figure 2 and Figure 4 The method in the corresponding embodiment will therefore not be described in detail here.
[0183] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, so that the computer device can perform the above-mentioned Figure 2 and Figure 4 The method in the corresponding embodiment will therefore not be described in detail here.
[0184] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0185] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present application are still within the scope covered by the present application.
Claims
1. A data processing method, characterized in that: The method comprises: During real-time communication between a first device and a second device, receiving an audio data packet transmitted by the first device; the audio data packet includes an audio signal on the first device side and first delay information, the first delay information being used to indicate delays incurred at various transmission processing stages experienced by the first device side during transmission of the audio signal along the audio link; performing delay statistical processing on the audio data packet to obtain second delay information of the audio signal; the second delay information is used to indicate delays incurred at various receiving and processing stages experienced by the audio signal on the second device during transmission along the audio link; The first delay information of the audio signal and the second delay information of the audio signal are integrated to obtain a target audio delay.
2. The method according to claim 1, wherein The performing delay statistical processing on the audio data packet to obtain second delay information of the audio signal includes: performing delay statistical processing at each receiving and processing stage of the audio link according to the audio data packet to obtain a delay of the audio signal generated at each receiving and processing stage; Delays generated by the audio signal in various receiving and processing stages are integrated to obtain second delay information of the audio signal.
3. The method according to claim 2, wherein The receiving and processing stage includes an unpacking stage, the audio data packet also includes a packing timestamp, and performing delay statistical processing at each receiving and processing stage of the audio link according to the audio data packet to obtain the delay generated by the audio signal at each receiving and processing stage, including: In the unpacking stage, the audio data packet is unpacked to obtain a packing timestamp in the audio data packet; Determining a depacketization timestamp of the audio data packet, and obtaining a network transmission delay of the audio signal according to the packing timestamp and the depacketization timestamp of the audio data packet; The network transmission delay is determined as the delay generated by the audio signal in the depacketization stage.
4. The method according to claim 2, wherein The receiving and processing stage includes a decoding stage, and performing delay statistical processing at each receiving and processing stage of the audio link according to the audio data packet to obtain the delay generated by the audio signal at each receiving and processing stage includes: Unpacking the audio data packet to obtain an audio stream corresponding to the audio data packet; In the decoding stage, delay statistics processing is performed on the audio stream obtained by the unpacking process to obtain the decoding delay of the audio signal; The decoding delay is determined as the delay generated by the audio signal in the decoding stage.
5. The method according to claim 4, wherein In the decoding stage, performing delay statistics processing on the audio stream obtained by depacketization to obtain the decoding delay of the audio signal includes: During the decoding phase, decoding the audio stream, counting a duration spent on caching decoding reference data generated during the decoding process, and after the decoding process is completed, determining the duration spent on caching the decoding reference data generated during the decoding process as a decoding cache delay of the audio signal; performing a cache process on the audio data obtained by the decoding process, counting a duration of the cache process during the cache process, and determining a data cache delay of the audio signal based on the duration of the cache process; The data buffer delay of the audio signal and the decoding buffer delay of the audio signal are integrated to obtain the decoding delay of the audio signal.
6. The method according to claim 2, wherein The receiving and processing stage includes a playing stage, and performing delay statistical processing at each receiving and processing stage of the audio link according to the audio data packet to obtain the delay generated by the audio signal at each receiving and processing stage includes: Decoding the audio data packet to obtain audio data corresponding to the audio data packet; During the playback phase, performing delay statistical processing on the audio data to obtain a playback delay of the audio signal; The playback delay is determined as the delay generated by the audio signal during the playback phase.
7. The method according to claim 6, wherein The number of the audio data packet includes one, and the decoding process obtains one audio data; and in the playback phase, performing delay statistical processing according to the audio data to obtain the playback delay of the audio signal, including: During the playback phase, performing digital-to-analog conversion on the audio data, and during the digital-to-analog conversion, counting the duration of the digital-to-analog conversion to obtain a second hardware delay of the audio signal; performing a buffering process on the audio signal, and during the buffering process, counting a duration spent on buffering the audio signal, and determining the duration spent on buffering the audio signal as a playback buffering delay of the audio signal; The playback buffer delay of the audio signal and the second hardware delay of the audio signal are integrated to obtain the playback delay of the audio signal.
8. The method according to claim 6, wherein The number of the audio data packets includes a plurality, and the plurality of audio data packets includes audio data packets transmitted by the first device and audio data packets transmitted by other devices; the decoding process obtains a plurality of audio data, and the receiving process stage includes a stream mixing stage; and in the playback stage, before performing delay statistical processing based on the audio data to obtain the playback delay of the audio signal, the method further includes: In the mixing stage, the plurality of audio data are mixed to obtain mixed stream data; the mixed stream data is used to determine the playback delay of the audio signal in the playback stage; During the stream mixing process, counting the time spent on the stream mixing process, and after the stream mixing process is completed, determining the time spent on the stream mixing process as the stream mixing delay of each audio signal; The mixing delay of each audio signal is determined as the delay generated by the corresponding audio signal in the mixing stage.
9. The method according to claim 2, wherein The integrating the delays generated by the audio signal in each receiving and processing stage to obtain second delay information of the audio signal includes: The delays generated by the audio signal in each receiving and processing stage are combined to obtain a first delay set, and the first delay set is determined as the second delay information of the audio signal.
10. The method according to claim 1, wherein The transmission processing stage includes one or more of the following: an acquisition stage, a preprocessing stage, and an encoding stage, and the first delay information includes: an acquisition delay generated in the acquisition stage, a preprocessing delay generated in the preprocessing stage, and an encoding delay generated in the encoding stage; The acquisition delay is composed of an acquisition buffer delay and a first hardware delay, wherein the first hardware delay refers to the time it takes for the first device to perform analog-to-digital conversion on the collected audio signal during the acquisition phase; the acquisition buffer delay refers to the time it takes for the first device to cache the audio data obtained by the analog-to-digital conversion during the acquisition phase; The preprocessing delay refers to the duration of the buffering operation taken by the first device during the preprocessing phase to preprocess the audio data; The encoding delay includes: during the encoding process of the pre-processed audio data in the encoding stage, the length of time it takes to cache the reference data generated in the encoding process.
11. The method according to any one of claims 1 to 10, wherein The integrating the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay includes: combining the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay; or The first delay information of the audio signal and the second delay information of the audio signal are summed to obtain a target audio delay.
12. The method according to any one of claims 1 to 10, wherein The first device and the second device communicate in real time via a target application, and the audio data packet is sent to the second device via an application server of the target application; after performing delay statistical processing on the audio data packet to obtain second delay information of the audio signal, the method further includes: The second delay information of the audio signal is sent to the application server of the target application, so that the application server performs delay troubleshooting on the audio link based on the second delay information and the first delay information.
13. The method according to claim 12, wherein: Each processing stage is used to indicate the corresponding processing logic for the audio signal. The type of any processing stage is a transmission processing stage or a reception processing stage. The delay troubleshooting includes at least one of the following: If the delay of the audio signal in any processing stage is greater than or equal to the first preset delay, the processing logic indicated by the corresponding processing stage is optimized; or, If the sum of the delays generated by the audio signal in each transmission processing stage is greater than or equal to a second preset delay, selecting the transmission processing stage with the largest delay from the transmission processing stages as the target processing stage, and optimizing the processing logic indicated by the target processing stage; or, If the sum of the delays generated by the audio signal in each receiving and processing stage is greater than or equal to a third preset delay, the receiving and processing stage with the largest delay is selected from the various receiving and processing stages as the target processing stage, and the processing logic indicated by the target processing stage is optimized.
14. The method according to any one of claims 1 to 10, wherein The first device and the second device perform real-time communication via a target application, the target audio delay is obtained by summing the first delay information of the audio signal and the second delay information of the audio signal, and the method further includes: If the target audio delay is greater than or equal to the preset audio delay, or the number of audio signals whose target audio delay is greater than or equal to the preset audio delay exceeds a preset threshold, output a delay prompt; The delay prompt is used to prompt adjustment of the operation settings of the target application; the operation settings include at least one of the following: the frequency of collecting audio signals; the rate of transmitting audio signals.
15. The method according to any one of claims 1 to 10, wherein The real-time communication between the first device and the second device includes any one of the following: real-time communication in a live broadcast scenario, real-time communication in an audio call scenario, real-time communication in a video call scenario, real-time communication in a chorus scenario, real-time communication in a game scenario, and real-time communication in an online meeting scenario.
16. A data processing device, characterized in that: include: a receiving unit, configured to receive an audio data packet transmitted by a first device during real-time communication between the first device and the second device; The audio data packet includes an audio signal on the first device side and first delay information, where the first delay information is used to indicate delays incurred at various transmission processing stages experienced by the audio signal on the first device side during transmission along the audio link. a processing unit, configured to perform delay statistical processing on the audio data packet to obtain second delay information of the audio signal; The second delay information is used to indicate: delays generated by each receiving and processing stage experienced by the audio signal on the second device during transmission according to the audio link; The processing unit is further configured to integrate the first delay information of the audio signal and the second delay information of the audio signal to obtain a target audio delay.
17. A computer device, characterized in that: include: a processor suitable for executing a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the data processing method according to any one of claims 1 to 15 is executed.
18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data processing method according to any one of claims 1 to 15 is executed.
19. A computer program product, characterized in that The computer program product includes a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the data processing method according to any one of claims 1 to 15.