Audio data processing method and related device
By identifying the audio frame type and adjusting the audio frame duration in combination with the buffer delay, the problem of discontinuous playback of audio data caused by network jitter in real-time communication is solved, and the user's auditory experience is improved.
Patent Information
- Application Number
- CN202410101381.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-24
- Publication Date
- 2025-07-25
AI Technical Summary
In real-time communication, due to network jitter, the prior art cannot fully compensate for the buffer buffer delay, resulting in distortion of audio data and affecting the user's auditory experience.
By identifying the type of audio frame and adjusting the duration of the audio frame in combination with the buffer delay, such as extending or shortening the playback time of music frames and non-vocal frames, to achieve continuous playback of audio data and avoid adjusting the duration of vocal frames.
It realizes continuous playback of audio data, improves the user's auditory experience, and avoids audio distortion.
Smart Images

Figure CN120378416A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, and in particular, to an audio data processing method and related devices. Background Art
[0002] With the development of Internet technologies and communication technologies, voice calls based on voice packet switching are increasingly favored by users. Taking real-time communication (RTC) as an example, during real-time communication, due to network jitter, a buffer is usually set at the receiving end to mitigate the impact of network jitter. However, since the network jitter in real-time communication is random, that is, there are cases where the buffer size of the jitter buffer cannot meet the network jitter. During the process of processing the received audio data, if the network jitter corresponding to the audio data is greater than the buffer delay, in order to achieve continuous playback of the audio data, the audio data to be played can be slow-played (i.e., the playback speed is slowed down); if the network jitter corresponding to the audio data is less than the buffer delay, in order to achieve continuous playback of the audio data, the audio data to be played can be fast-played (i.e., the playback speed is increased). Fast-playing or slow-playing the audio data both changes the actual duration of the audio data, resulting in audio data distortion and seriously affecting the user's auditory experience. Summary of the Invention
[0003] This application provides an audio data processing method and related devices. After receiving an audio frame of audio data and storing the audio frame in a buffer, when the network jitter exceeds the buffer delay range of the buffer or the sound card jitter exceeds the buffer delay range of the buffer, the duration of the audio frame with a specific type of frame content is adjusted to ensure the continuous playback effect of the audio data, and the duration of other types of audio frames is not adjusted, avoiding audio data distortion, thereby enhancing the user's auditory experience.
[0004] To achieve the above object, this application adopts the following technical solutions:
[0005] In a first aspect, an audio data processing method is provided. The method includes: receiving an audio frame of audio data and storing the audio frame in a buffer; if the network jitter corresponding to the audio frame exceeds the buffer delay range of the buffer or the sound card jitter exceeds the buffer delay range of the buffer, determining the type of the frame content of the audio frame; if the type of the frame content of the audio frame is a first type, adjusting the duration of the audio frame based on a first relationship, where the first relationship is the relationship between the network jitter or the sound card jitter and the buffer delay range.
[0006] Thus, after receiving an audio frame of audio data, the audio frame is stored in a buffer; if the network jitter or sound card jitter corresponding to the audio frame exceeds the buffer delay range of the buffer, the type of the frame content of the audio frame is determined; if the frame content of the audio frame is of the first type, the duration of the audio frame is adjusted according to the relationship between the network jitter or sound card jitter and the buffer delay, so that the receiving end plays the audio frame according to the adjusted duration; if the frame content of the audio frame is of other types, the duration of the audio frame is not adjusted, that is, the receiving end plays the audio frame according to the actual duration of the audio frame; by adjusting the duration of the audio frame whose frame content in the audio data is of the first type, continuous playback of the audio data is achieved; since the duration of the audio frames corresponding to other frame contents of the audio data remains unchanged, distortion of the audio of this audio frame is avoided, so as to improve the user's auditory experience.
[0007] In some embodiments, determining the type of the frame content of the audio frame includes: determining whether the audio frame is a music frame; if the audio frame is not a music frame, determining the type of the frame content of the audio frame.
[0008] Thus, before determining the type of the frame content of the audio frame, it can be detected whether the audio frame is a music frame. If the audio frame is not a music frame, that is, it does not have the characteristics of continuous playback of a music frame. Then, the type of the frame content of the audio frame is determined. If the type of the frame content of the audio frame is of the first type, the duration of the audio frame can be adjusted according to the relationship between the network jitter and the buffer delay range, so as to achieve continuous playback of the audio data; if the type of the frame content of the audio frame is not of the first type, the duration of the audio frame does not need to be adjusted, so as to play the audio frame according to the actual duration.
[0009] In some embodiments, determining the type of the frame content of the audio frame includes: obtaining the signal-to-noise ratio of the audio frame; if the signal-to-noise ratio is less than a first threshold, determining that the type of the frame content of the audio frame is of the first type.
[0010] Thus, since the playback sound of an audio frame with a low signal-to-noise ratio is low, that is, in a real-time communication process, it has little impact on the user's auditory experience. Therefore, the type of the frame content of the audio frame is determined by the relationship between the signal-to-noise ratio of the audio frame and the first threshold, so that by adjusting the playback duration of the audio frame with a low signal-to-noise ratio, continuous playback of the audio data is achieved, and the audio frame with a high signal-to-noise ratio can be played according to the normal playback duration, so as to improve the user's auditory experience.
[0011] In some embodiments, the method further includes: if the audio frame is a music frame, obtaining the signal-to-noise ratio of the audio frame; if the signal-to-noise ratio is less than a second threshold, determining that the type of the frame content of the audio frame is of the first type.
[0012] Thus, when the network jitter of an audio frame exceeds the buffer delay range of the buffer, and this audio frame is a music frame. Since music frames are continuous and the sound of music frames has ups and downs, and usually the playback sound with a low signal-to-noise ratio is lower. Therefore, among the audio data of the music type, the audio frames with a signal-to-noise ratio lower than the second threshold are adjusted in playback duration so as to continuously play the audio data. For the audio frames with a signal-to-noise ratio greater than or equal to the second threshold, they are played according to the original playback duration to enhance the user's auditory experience.
[0013] In some embodiments, the first type includes non-human voices. If the type of the frame content of the audio frame is the first type, then adjusting the duration of the audio frame based on the relationship between the network jitter and the buffer delay range includes: if the type of the frame content of the audio frame is non-human voice, then adjusting the duration of the audio frame based on the relationship between the network jitter and the buffer delay range.
[0014] Thus, the types of the frame content of the audio frames include human voices and non-human voices. After identifying the type of the frame content of the audio frame, if the type of the frame content of the audio frame is a human voice, then the duration of this audio frame is not adjusted, that is, this audio frame is played according to the actual duration; if the type of the frame content of the audio frame is non-human voice (i.e., the first type), then the duration of this audio frame is adjusted, that is, the duration of the audio frame is extended or shortened, and it is played according to the extended or shortened duration. Since most scenarios of real-time communication are multiple-user calls, the communication content is mainly the interactive transmission of the voices of multiple users. Since there are gaps during the communication interaction process of users, such as pauses or waiting between sentences. The type of the frame content of the audio frame corresponding to this gap is non-human voice; when the network jitter of the audio frame exceeds the buffer delay range of the buffer, by adjusting the playback duration of the audio frame corresponding to non-human voice, the continuous playback of the audio data is achieved; since the above process does not change the playback duration of the audio frame corresponding to human voice, that is, the part related to human voice in real-time communication can be played according to the actual duration of this audio frame. That is, the continuity of real-time communication is achieved, and the user's auditory experience can be enhanced.
[0015] In some embodiments, the method further includes: receiving an instruction message from the user, and determining the first type according to the instruction message.
[0016] In some embodiments, adjusting the duration of the audio frame based on the first relationship includes: if the network jitter or the sound card jitter is greater than the value within the buffer delay range, then extending the duration of the audio frame; if the network jitter or the sound card jitter is less than or equal to the value within the buffer delay range, then shortening the duration of the audio frame.
[0017] In a second aspect, there is provided an audio data processing apparatus, comprising: a memory including computer-readable instructions; and a processor communicatively coupled to the memory, the processor being configured to execute the computer-readable instructions such that the audio data processing apparatus performs the audio data processing method according to any one of the first aspect.
[0018] In a third aspect, there is provided a computer-readable storage medium including a program or instructions which, when executed by a processor, implement the audio data processing method according to any one of the first aspect.
[0019] In a fourth aspect, there is provided a chip including a processor configured to call and run instructions stored in a memory from the memory such that an audio data processing apparatus installed with the chip performs the audio data processing method according to any one of the first aspect.
[0020] For the beneficial effects brought by each possible implementation manner of the audio data processing apparatus provided in the second aspect, the computer-readable storage medium provided in the third aspect, and the chip provided in the fourth aspect of the embodiments of the present application, reference may be made to the descriptions in various possible implementation manners of the first aspect, and details are not described herein again. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 FIG. is a schematic diagram of a network architecture provided by an embodiment of the present application;
[0022] Figure 2 FIG. is a schematic flowchart of an audio data processing method provided by an embodiment of the present application;
[0023] Figure 3 FIG. is a schematic structural diagram of an audio data processing apparatus provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The technical solutions in the present application will be described below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments.
[0025] First, the technical terms related to the embodiments of the present application are introduced:
[0026] 1. Jitter Buffer
[0027] A jitter buffer is a shared data area in which voice packets are collected, stored, and sent to a voice processor at regular intervals. The variation in the arrival time of packets is called jitter, which is caused by network congestion, timing drift, or routing changes. The jitter buffer is placed at the receiving end of the voice connection, and it deliberately delays the arriving packets so that the end user will experience a clear connection without voice distortion.
[0028] Please refer to Figure 1 , Figure 1 , which is a schematic diagram of a network architecture provided by an embodiment of the present application. Figure 1 The network architecture includes: Terminal 1 and Terminal 2, and Terminal 1 communicates with Terminal 2 through a network. Among them, Terminal 1 can collect audio data through devices such as a microphone, and the audio data can be the user's voice or music, etc. And Terminal 1 can send the audio data to Terminal 2 in the form of audio frames in real time. Terminal 2 can receive the audio data and play the audio data through the sound card of the terminal.
[0029] It is easy to understand that during the process of transmitting audio data between Terminal 1 and Terminal 2, due to problems such as network congestion, insufficient bandwidth, and routing problems, the delays of different audio data transmitted through the same connection between Terminal 1 and Terminal 2 are different, that is, network jitter exists.
[0030] In order to alleviate the impact of network jitter on audio playback, a buffer can be set in Terminal 2 (for example, implemented through a jitter buffer). After receiving the audio data, Terminal 2 first stores the audio data in the buffer, and delays and compensates for network jitter through the buffer to avoid interruptions or discontinuities when the sound card plays audio. However, when the network jitter is large or small, the buffer delay of the buffer alone cannot achieve compensation for network jitter. In order to achieve continuous playback of audio data, the buffer delay of the buffer can be combined with the adjustment of the duration of audio frames to achieve compensation for network jitter, so that the audio data can be continuously played. For example, when the network jitter is greater than the buffer delay of the buffer, the audio data to be played can also be stretched (that is, the time of the audio data is extended); when the network jitter is less than the buffer delay of the buffer, the audio data to be played can be compressed (that is, the time of the audio data is reduced). By adjusting the duration of the audio data in combination with the buffer delay of the buffer, compensation for network jitter is achieved, so that the audio data is continuously played.
[0031] It is easy to understand Figure 1After the middle terminal 2 receives the audio frame, it stores the audio frame in the buffer. And the audio frame in the buffer is played through the sound card. Due to hardware problems of the sound card itself, sound card driver problems, sound card settings and other problems, the playback beat of the sound card of terminal 2 jitters frequently. And after the jitter of the sound card exceeds the buffer delay range of the buffer, however, when the jitter of the sound card is large or small, it is impossible to compensate for the jitter of the sound card only through the buffer delay of the buffer. In order to achieve continuous playback of audio data, the buffer delay of the buffer and the adjustment of the duration of the audio frame can be combined to compensate for the jitter of the sound card, so that the audio data can be continuously played. For example, when the jitter of the sound card is greater than the buffer delay of the buffer, stretching processing can also be performed on the audio data to be played (that is, extending the time of the audio data); when the jitter of the sound card is less than the buffer delay of the buffer, compression processing can be performed on the audio data to be played (that is, reducing the time of the audio data). By adjusting the duration of the audio data in combination with the buffer delay of the buffer, the jitter of the sound card is compensated, so that the audio data is continuously played.
[0032] However, in the above scenario, since the duration of the audio data is adjusted, that is, the playback duration of the audio data does not match the actual duration of the audio data, the voice played by terminal 2 is very different from the original sound of the audio data, affecting the user experience.
[0033] Based on the above problems, the embodiment of the present application provides an audio data processing method. After receiving the audio frame of the audio data, the audio frame is stored in the buffer; if the network jitter or sound card jitter corresponding to the audio frame exceeds the buffer delay range of the buffer, the type of the frame content of the audio frame is determined; if the frame content of the audio frame is of the first type, the duration of the audio frame is adjusted according to the relationship between the network jitter or sound card jitter and the buffer delay, so that the receiving end plays the audio frame according to the adjusted duration; if the frame content of the audio frame is of other types (that is, other types except the first type), the duration of the audio frame is not adjusted, that is, the receiving end plays the audio frame according to the actual duration of the audio frame; by adjusting the duration of the audio frame whose frame content type in the audio data is of the first type, continuous playback of the audio data is achieved; since the duration of the audio frame corresponding to other frame contents of the audio data remains unchanged, the audio of the audio frame is prevented from being distorted, so as to improve the auditory experience of the user.
[0034] Figure 1 The middle terminal 1 and the terminal 2 can be connected through a network, and the network can be a wireless network or a wired network. Or, the terminal 1 and the terminal 2 can achieve audio transmission through the client of an application program including an instant messaging application.
[0035] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. Additionally, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or proprietary data communication technologies can be used to replace or supplement the above data communication technologies.
[0036] Both terminal 1 and terminal 2 can be various electronic devices, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, wearable devices, intelligent voice interaction devices, vehicle-mounted terminals, smart home appliances, aircraft, augmented reality devices, virtual reality devices, etc.
[0037] Optionally, the clients of the application programs installed in terminal 1 and terminal 2 are the same, or clients of the same type of application program based on different operating systems. Depending on the different terminal platforms, the specific form of the client of the application program can also be different. For example, the client of the application program can be a mobile phone client, a PC client, etc.
[0038] Figure 1 The numbers of terminal 1 and terminal 2 in [description] are illustrative. According to actual needs, there can be any number of terminal 1 and terminal 2. The embodiments of the present disclosure do not limit this.
[0039] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an audio data processing method provided by an embodiment of this application. Figure 2The method shown is described by taking the example of network jitter exceeding the buffer delay range of the buffer. It can be understood that the above embodiments are also equally applicable to the scenario where the sound card jitter exceeds the buffer delay of the buffer. Figure 2 The audio data processing method in
[0040] S201: Receive the audio frames of the audio data and store the audio frames in the buffer.
[0041] It is easy to understand that since the real-time communication network stores network jitter, the network jitter affects the communication quality of real-time communication. To reduce the impact of network jitter on real-time communication, a buffer can be set at the receiving end of the audio data. After the receiving end receives the audio data, the audio data is stored in the buffer to compensate for the network jitter of the audio data and improve the communication quality of real-time communication.
[0042] S202: If the network jitter of the audio frame exceeds the buffer delay range of the buffer, determine the type of the frame content of the audio frame.
[0043] It is easy to understand that since the network jitter of the audio frame exceeds the buffer delay range of the buffer, that is, the network jitter cannot be successfully compensated only by the buffer delay of the buffer. To achieve continuous playback of audio data, the buffer delay of the buffer and other processing of the audio frame can be combined. For example, extend the duration of each audio frame or shorten the duration of each audio frame.
[0044] Optionally, the type of the frame content of the audio frame can be identified by voice classification or voice recognition technology. For example, identify whether the current audio frame is human voice, dog barking, bird singing, etc. For example, a deep neural network model can be used to identify various sounds.
[0045] Optionally, during a real-time call, the audio data transmitted mainly includes human voice and non-human voice, then the type of the frame content of the current audio frame can be identified by voice activity detection. Determine whether the frame content type of the current frame is human voice or non-human voice.
[0046] Optionally, voice activity detection can be implemented by deep learning technology, machine learning algorithms, endpoint detection algorithms, etc.
[0047] Optionally, the buffer delay range can be determined by the buffer delay of the buffer and the safety range. Exemplarily, the buffer delay is 8 milliseconds, and the safety range is ±2 milliseconds, then the buffer delay range is: 6 milliseconds to 10 milliseconds.
[0048] Optionally, the network jitter corresponding to the audio frame can be the network jitter of the network when the receiving end device receives the audio frame.
[0049] S203: If the type of the frame content of the audio frame is the first type, adjust the duration of the audio frame according to the relationship between the network jitter and the buffer delay range.
[0050] Optionally, after determining the type of the frame content of the audio frame, it is determined whether the type is the first type; if the type of the frame content of the audio frame is the first type, the duration of the audio frame is adjusted, and the adjusted duration of the audio frame and the buffer delay of the buffer are combined so that the audio data can be continuously played; if the frame content of the audio frame is not the first type, the playback duration of the audio frame is not adjusted, that is, it is played according to the actual duration of the audio frame.
[0051] Optionally, the playback duration of the audio frame is adjusted according to the relationship between the network jitter and the buffer delay range: if the network jitter is greater than the value within the buffer delay range, the duration of the audio frame can be extended; if the network jitter is less than the value within the buffer delay range, the duration of the audio frame is shortened.
[0052] It is easy to understand that the first type can be non-human voice or other user-specified sound types, such as barking. Then, if the frame content of the audio frame is barking, when the network jitter duration of the audio frame exceeds the buffer delay range, the duration of the audio frame with the frame content type of barking can be extended.
[0053] Optionally, the adjustment of the duration of the audio frame can be achieved through the Time-Scale Modification (TSM) technology. TSM is an audio processing technology used to change the time scale of an audio signal. Specifically, it can accelerate or decelerate the playback speed of an audio signal while keeping the pitch of the audio signal unchanged.
[0054] In this way, during a real-time communication process, after the receiving end receives an audio frame, the audio frame is stored in the buffer; if the network jitter of the audio frame exceeds the buffer range of the buffer. Since the network jitter of the audio frame cannot be compensated by the buffer delay of the buffer, it may cause the audio data to not be continuously played. By determining the type of the frame content of the audio frame and adjusting the duration of the audio frame of a specific type, and combining the buffer delay of the buffer to compensate for the network jitter, the audio data can be continuously played. And the playback duration of other types of audio frames is not adjusted, avoiding affecting the user's auditory experience due to adjusting the playback duration of other audio frames.
[0055] In some embodiments, the types of the frame content of an audio frame include human voice and non - human voice. After identifying the frame content of the audio frame, if the type of the frame content of the audio frame is human voice, the duration of the audio frame is not adjusted, that is, the audio frame is played according to the actual duration; if the type of the frame content of the audio frame is non - human voice (i.e., the first type), the duration of the audio frame is adjusted, that is, the duration of the audio frame is extended or shortened, and the audio frame is played according to the extended or shortened duration. Since most scenarios of real - time communication are multiple - user calls, the communication content is mainly the interactive transmission of the voices of multiple users. Since there are gaps during the communication interaction process of users, such as pauses or waiting between sentences and other scenarios. The type of the frame content of the audio frame corresponding to this gap is non - human voice; when the network jitter of the audio frame exceeds the buffer delay range of the buffer, the playback duration of the audio frame corresponding to non - human voice is adjusted to achieve continuous playback of audio data; since the above process does not change the playback duration of the audio frame corresponding to human voice, that is, the human - voice - related part in real - time communication can be played according to the actual duration of the audio frame. That is, the continuity of real - time communication is achieved, and the auditory experience of users can be improved.
[0056] In some embodiments, the first type includes multiple subtypes. For example, if the types of the frame content of an audio frame include human voice and non - human voice, then the non - human voice can be silence, bird calls, car horn sounds, etc. Then, if the network jitter of the audio frame exceeds the buffer delay range of the buffer and the type of the frame content of the audio frame is any one of multiple non - human voices, the duration of the audio frame can be adjusted; if the frame content of the audio frame is human voice, the duration of the audio frame does not need to be adjusted.
[0057] In some embodiments, if the types of the frame content of an audio frame include the first type and the second type, if the network jitter of the audio frame exceeds the buffer range of the buffer and the type of the frame content of the audio frame is the first type, the playback duration of the audio frame is adjusted; if the type of the frame content of the audio frame is the second type, the playback duration of the audio frame is not adjusted. Here, the second type can be non - the first type, or it can be one of the types of other multiple frame contents.
[0058] It is easy to understand that if the types of the frame content of an audio frame include multiple types, when the network jitter of the audio frame exceeds the buffer range of the buffer, the types that need to adjust the length of the audio frame and the types that do not need to adjust the length of the audio frame can be selected according to the actual needs of the user from the multiple types to achieve adaptive adjustment. Before S203, the method further includes: receiving indication information of the user, and determining a first type according to the indication information. The user can input the indication information by means such as pressing a key or touching, and the indication information carries the type of the audio frame whose length can be adjusted, so as to determine the first type according to the indication information of the user. When the network jitter of the audio frame exceeds the buffer range of the buffer, the playback duration of the audio frame whose frame content type is the first type in the audio data can be adjusted, so that the receiving end can continuously play the audio data.
[0059] It is easy to understand that since the signal-to-noise ratio of an audio frame affects the playback sound of the audio frame. If the signal-to-noise ratio of the audio frame is low, the clarity and quality of the audio are low, and the playback sound is small; if the signal-to-noise ratio is high, the clarity and quality of the audio are high, and the playback sound is large. Therefore, in S202, when the network jitter of the audio frame exceeds the buffer range of the buffer, the type of the frame content of the audio frame whose signal-to-noise ratio is lower than the first threshold can be used as the first type, and the playback duration of the audio frame can be adjusted; the type of the frame content of the audio frame whose signal-to-noise ratio is greater than or equal to the first threshold can be used as other types and the playback duration of the audio frame is not required, that is, the type of the frame content of the audio frame is determined by the relationship between the signal-to-noise ratio of the audio frame and the first threshold. Since the playback sound of the audio frame with a low signal-to-noise ratio is low, that is, in the process of real-time communication, the impact on the user's auditory experience is small. Therefore, the type of the frame content of the audio frame is determined by the relationship between the signal-to-noise ratio of the audio frame and the first threshold, so that by adjusting the playback duration of the audio frame with a low signal-to-noise ratio, the continuous playback of the audio data can be realized, and the audio frame with a high signal-to-noise ratio can be played according to the normal playback duration, so as to improve the user's auditory experience.
[0060] It is easy to understand that in a real-time communication scenario, the two ends of the real-time communication can be two users who are conducting a voice communication, and at this time, the communication voice of the user for audio data transmission. The two ends of the real-time communication can also be other situations, for example, one end of the real-time communication sends music voice to the other end. Music data has continuity. Because music is a continuous audio stream and needs to maintain continuous transmission. For a voice call, due to the characteristics of the speaker, the pauses between speaking or the pauses between sentences, etc., the call data has discontinuity.
[0061] In some embodiments, in S202, if the network jitter of the audio frame exceeds the buffer delay range of the buffer, before determining the type of the frame content of the audio frame, it is possible to detect whether the audio frame is a music frame. If the audio frame is a non-music frame, that is, it does not have the characteristics of continuous playback of a music frame. Then determine the type of the frame content of the audio frame. If the type of the frame content of the audio frame is the first type, the duration of the audio frame can be adjusted according to the relationship between the network jitter and the buffer delay range, so as to continuously play the audio data; if the type of the frame content of the audio frame is not the first type, the duration of the audio frame does not need to be adjusted, so as to play the audio frame according to the actual duration.
[0062] In some embodiments, in S202, if the network jitter of the audio frame exceeds the buffer delay range of the buffer, before determining the type of the frame content of the audio frame, it is possible to detect whether the audio frame is a music frame. If the audio frame is a music frame, obtain the signal-to-noise ratio of the audio frame. If the signal-to-noise ratio is lower than the second threshold, adjust the playback duration of the audio frame; if the signal-to-noise ratio is greater than or equal to the second threshold, do not adjust the playback duration of the audio frame. In this way, when the network jitter of the audio frame exceeds the buffer delay range of the buffer and the audio frame is a music frame. Since music frames have continuity and the sound of music frames has ups and downs, and usually the playback sound with a low signal-to-noise ratio is lower. Therefore, in the audio data of the music type, the playback duration of the audio frames with a signal-to-noise ratio lower than the second threshold is adjusted, so as to continuously play the audio data. The audio frames with a signal-to-noise ratio greater than or equal to the second threshold are played according to the original playback duration to improve the user's auditory experience.
[0063] Optionally, a deep learning method can be used to detect whether the audio frame is a music frame. For example, a deep learning model is constructed using models such as a deep neural network (DNN) or a convolutional neural network (CNN), and the trained deep learning model is used to identify and detect whether the current audio frame is a music frame. Of course, in other embodiments, the audio frame can also be detected by other means to detect whether the audio frame is a music frame.
[0064] In some embodiments, in order to improve the accuracy of the detection of the current speech frame, multiple audio rows in the buffer can be detected to determine that the current scene is a music scene. For example, if the multiple speech frames transmitted are all music-related speech, it can be determined that the scene corresponding to the multiple speech frames is a music scene, and then the multiple speech frames can be processed according to the music frames in the music scene. For example, if it is determined that the current scene is a music scene based on multiple audio frames, then the multiple audio frames are music frames.
[0065] In some embodiments, after the receiving end receives the audio frame, the audio frame is stored in the buffer. If the network jitter corresponding to the audio frame or the sound card jitter corresponding to the audio frame exceeds the buffer delay range of the buffer, then detect whether the audio frame is a music frame;
[0066] If the audio frame is a non-music frame, then detect whether the type of the frame content of the audio frame is the first type; if the type of the frame content of the audio frame is the first type and the network jitter or sound card jitter is greater than the value within the buffer delay range, then extend the duration of the audio frame; if the type of the frame content of the audio frame is the first type and the network jitter or sound card jitter is less than or equal to the value within the buffer delay range, then shorten the duration of the audio frame; if the type of the frame content of the audio frame is not the first type, then the duration of the audio frame needs to be adjusted.
[0067] If the audio frame is a music frame, then obtain the signal-to-noise ratio of the audio frame. If the signal-to-noise ratio of the audio frame is greater than the preset threshold, then there is no need to adjust the duration of the audio frame; if the signal-to-noise ratio of the audio frame is less than or equal to the preset threshold, and the network jitter is greater than the value within the buffer delay range or the sound card jitter is greater than the value within the buffer delay range, then extend the duration of the audio frame; if the network jitter is less than or equal to the value within the buffer delay range or the sound card jitter is less than or equal to the value within the buffer delay range, then shorten the duration of the audio frame.
[0068] It should be understood that the above is only to help those skilled in the art better understand the embodiments of the present application, rather than to limit the scope of the embodiments of the present application. Those skilled in the art can obviously make various equivalent modifications or changes according to the above examples. For example, some steps in the above methods may not be necessary, or some steps may be newly added, etc. Or any combination of any two or any multiple of the above embodiments. Such modified, changed or combined solutions also fall within the scope of the embodiments of the present application.
[0069] It should also be understood that the classification of the manners, situations, categories and embodiments in the embodiments of the present application is only for the convenience of description and should not constitute a special limitation. The features in various manners, categories, situations and embodiments can be combined without conflict.
[0070] It should also be understood that the various digital numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. The magnitudes of the serial numbers of the above processes do not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0071] It should also be understood that the above description of the embodiments of the present application focuses on emphasizing the differences between the embodiments. The same or similar parts not mentioned can be referred to each other. For the sake of brevity, they will not be elaborated here.
[0072] The above combination Figure 1 - Figure 2 describes the embodiments of the method and system provided by the embodiments of the present application. Next, an audio data processing device provided by the embodiments of the present application will be described.
[0073] In this embodiment, the audio data processing device can be divided into functional modules according to the above method. For example, it can be divided into various functional modules corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0074] It should be noted that the relevant content of each step involved in the above method embodiment can be cited to the function description of the corresponding functional module, and will not be repeated here.
[0075] The audio data processing device provided by the embodiment of the present application is used to execute the audio data processing method provided by the above method embodiment, so it can achieve the same effect as the above implementation method.
[0076] In other embodiments, in the case of adopting an integrated unit, the audio data processing device may include a processing module, a storage module, and a communication module. Among them, the processing module can be used to control and manage the actions of the audio data processing device. For example, it can be used to support the audio data processing device to execute the steps performed by the processing unit. The storage module can be used to support the storage of program codes and data, etc. The communication module can be used to support the communication between the audio data processing device and other network devices and audio data processing devices.
[0077] Among them, the processing module can be a processor or a controller. It can be used to implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. The processor can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, and so on. The storage module can be a memory. The communication module can specifically be a device for interacting with other audio data processing devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, etc.
[0078] Based on the same concept, the embodiment of the present application also provides an audio data processing device. Refer to Figure 3 , Figure 3 which shows a schematic structural diagram of an exemplary audio data processing device of the present application. Figure 3 The audio data processing device shown can execute the steps in any audio data processing method executed by the audio data processing device provided by the embodiment of the present application.
[0079] The audio data processing device 300 includes at least one processor 301, a memory 303, and at least one network interface 304.
[0080] The processor 301 is, for example, a general-purpose CPU, a digital signal processor (DSP), a network processor (NP), a GPU, a neural network processing unit (NPU), a data processing unit (DPU), a microprocessor, or one or more integrated circuits or application-specific integrated circuits (ASICs) for implementing the solution of this application, a programmable logic device (PLD), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The PLD is, for example, a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It can implement or execute various logic blocks, modules, and circuits described in connection with the disclosure of this application. The processor can also be a combination for implementing computing functions, such as including a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on.
[0081] Optionally, the audio data processing device 300 further includes a bus 302. The bus 302 is used to transfer information between the components of the audio data processing device 300. The bus 302 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 302 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0082] The memory 303 is, for example, a read only memory (ROM) or other type of storage device that can store static information and instructions, such as a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, such as an electrically erasable programmable read only memory (EEPROM), a compact disc read only memory (CD ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 303 exists independently, for example, and is connected to the processor 301 via the bus 302. The memory 303 can also be integrated with the processor 301.
[0083] The network interface 304 uses any device such as a transceiver to communicate with other devices or communication networks, and the communication network can be an Ethernet, a radio access network (RAN) or a wireless local area network (WLAN), etc. The network interface 304 can include a wired network interface and can also include a wireless network interface. Specifically, the network interface 304 can be an Ethernet interface, such as a fast Ethernet (FE) interface, a gigabit Ethernet (GE) interface, an asynchronous transfer mode (ATM) interface, a WLAN interface, a cellular network interface, or a combination thereof. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. In some embodiments of the present application, the network interface 304 can be used for the audio data processing device 300 to communicate with other devices.
[0084] In a specific implementation, as some embodiments, the processor 301 can include one or more CPUs. Each of these processors can be a single-core processor or a multi-core processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0085] In specific implementations, as some embodiments, the audio data processing device 300 may include multiple processors. Each of these processors may be a single-core processor or a multi-core processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0086] In some embodiments, the memory 303 is used to store program instructions for executing the solution of this application, and the processor 301 may execute the program instructions stored in the memory 303. That is, the audio data processing device 300 may implement the method provided by the method embodiment shown in the above embodiment through the program instructions in the processor 301 and the memory 303. The program instructions may include one or more software modules. Optionally, the processor 301 itself may also store program instructions for executing the solution of this application.
[0087] In the specific implementation process, the processor 301 in the audio data processing device 300 of this application reads the instructions in the memory 303, so that Figure 3 the audio data processing device 300 shown can execute all or part of the steps in the audio data processing method executed by the audio data processing device in the above embodiment.
[0088] Among them, each step of the method described in the above embodiment is completed by the integrated logic circuit of the hardware in the processor of the audio data processing device 300 or the instructions in software form. Combining the steps of the method embodiment disclosed in this application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method embodiment. To avoid repetition, it will not be described in detail here.
[0089] It should be understood that the above-mentioned processor can be a central processing unit (CPU), or it can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It is worth noting that the processor can be a processor that supports the advanced RISC machines (ARM) architecture.
[0090] Furthermore, in an alternative embodiment, the above-mentioned memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. The memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0091] The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus random access memory (DR RAM).
[0092] The audio data processing device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0093] This application embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the above method embodiment is implemented.
[0094] This application embodiment also provides a computer program product. When the computer program product runs on an audio data processing device, the audio data processing device is caused to execute the method described in the above method embodiment.
[0095] This application embodiment provides a chip, including a processor, which is used to call and run instructions stored in a memory, so that a communication device installed with the chip executes the method described in any of the audio data processing devices provided in this application embodiment.
[0096] This application embodiment also provides a chip system, including a processor, where the processor is coupled to a memory, and the processor executes a computer program stored in the memory to implement the method described in the above method embodiment. Wherein, the chip system can be a single chip or a chip module composed of multiple chips.
[0097] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application embodiment are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0098] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium can include various media that can store program codes, such as ROM or random access memory RAM, magnetic disks, or optical discs.
[0099] In this application, the naming or numbering of steps does not mean that the steps in the method process must be executed in the time / logical order indicated by the naming or numbering. The named or numbered process steps can be changed in the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0100] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0101] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the apparatus or unit can be in an electrical, mechanical or other form.
[0102] It should be understood that in the description of this application specification and the appended claims, the terms "include", "comprise", "have" and any variations thereof are intended to cover non-exclusive inclusion, all meaning "including but not limited to", unless otherwise specifically emphasized in other ways. For example, a process, method, system, product or device that includes a series of steps or modules does not have to be limited to those steps or modules clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices.
[0103] In the description of this application, unless otherwise specified, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B can mean A or B; the "and / or" in this application is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural.
[0104] Also, in the description of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or plural.
[0105] As used in the specification and the appended claims of the present application, the term "if" may be construed as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0106] In addition, in the description of the specification and the appended claims of the present application, terms such as "first", "second", etc. are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that shown or described here; the features defined with "first" and "second" may explicitly or implicitly include at least one of such features.
[0107] In the embodiments of the present application, words such as "exemplarily" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.
[0108] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An audio data processing method, characterized in that The method includes: Receiving an audio frame of audio data and storing the audio frame in a buffer; If the network jitter corresponding to the audio frame exceeds the buffer delay range of the buffer or the sound card jitter exceeds the buffer delay range of the buffer, then determining the type of the frame content of the audio frame; If the type of the frame content of the audio frame is the first type, adjusting the duration of the audio frame based on a first relationship, where the first relationship is the relationship between the network jitter or the sound card jitter and the buffer delay range.
2. The method according to claim 1, characterized in that, The determining the type of the frame content of the audio frame includes: Judging whether the audio frame is a music frame; If the audio frame is not a music frame, determining the type of the frame content of the audio frame.
3. The method according to claim 2, characterized in that, The determining the type of the frame content of the audio frame includes: Obtaining the signal-to-noise ratio of the audio frame; If the signal-to-noise ratio is less than a first threshold, determining that the type of the frame content of the audio frame is the first type.
4. The method according to claim 2 or 3, characterized in that, The method further includes: If the audio frame is a music frame, obtaining the signal-to-noise ratio of the audio frame; If the signal-to-noise ratio is less than a second threshold, determining that the type of the frame content of the audio frame is the first type.
5. The method according to claim 1, wherein The first type includes non-human voices. If the type of the frame content of the audio frame is the first type, then adjusting the duration of the audio frame based on the relationship between the network jitter and the buffer delay range includes: If the type of the frame content of the audio frame is non-human voices, adjusting the duration of the audio frame based on the relationship between the network jitter and the buffer delay range.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Receiving an instruction message from a user and determining the first type according to the instruction message.
7. The method according to any one of claims 1 to 6, characterized in that, The adjusting the duration of the audio frame based on the first relationship includes: If the network jitter or the sound card jitter is greater than a value within the buffer delay range, extending the duration of the audio frame; If the network jitter or the sound card jitter is less than or equal to a value within the buffer delay range, shortening the duration of the audio frame.
8. An audio data processing device, characterized in that, Includes: A memory, where the memory includes computer-readable instructions; A processor communicating with the memory, where the processor is configured to execute the computer-readable instructions so that the audio data processing device executes the audio data processing method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, Includes a program or instructions which, when executed by a processor, implement the audio data processing method according to any one of claims 1-7.
10. A chip, characterized in that, Includes a processor configured to call and run instructions stored in a memory from the memory, so that an audio data processing device installed with the chip executes the audio data processing method according to any one of claims 1-7.