Audio packet loss data recovery methods, devices, electronic equipment and storage media

By aligning the target audio and the reference audio, the first reference frame and adjacent frames of the lost packet frame are obtained. Data recovery is performed based on audio feature association, which solves the problem of difficulty in balancing accuracy and latency in audio packet loss data recovery in existing technologies, and realizes efficient packet loss data recovery in multi-person online chorus applications.

CN117672234BActive Publication Date: 2026-07-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-08-25
Publication Date
2026-07-17

Smart Images

  • Figure CN117672234B_ABST
    Figure CN117672234B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, electronic device, and storage medium for audio packet loss data recovery. The method includes: aligning a target audio and a reference audio of the target audio to obtain an audio alignment result, wherein the target audio contains packet loss frames with missing data; obtaining a first reference frame in the reference audio corresponding to the position of the packet loss frame based on the position of the packet loss frame in the target audio and the audio alignment result, and obtaining a second reference frame located before and adjacent to the first reference frame; and performing data recovery on the packet loss frame based on the audio feature association between the first reference frame and the second reference frame, using the preceding target frame located before and adjacent to the packet loss frame as a reference. The embodiments of this application can simultaneously meet the accuracy requirements and low latency requirements of audio packet loss data recovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio, specifically to a method, apparatus, electronic device, and storage medium for recovering audio packet loss data. Background Technology

[0002] Since packet loss is inevitable during network data transmission, audio data is also inevitably subject to packet loss during network transmission. In various audio and video-related applications (such as online conferencing applications and online karaoke applications), audio packet loss can lead to audio stuttering and distortion. Therefore, to minimize audio stuttering and distortion, it is necessary to recover audio packet loss data. Existing audio packet loss data recovery technologies struggle to meet both accuracy and low latency requirements, and vice versa. Summary of the Invention

[0003] One objective of this application is to provide a method, apparatus, electronic device, and storage medium for audio packet loss data recovery that can simultaneously meet the accuracy and low latency requirements for audio packet loss data recovery.

[0004] According to one aspect of the embodiments of this application, an audio packet loss data recovery method is disclosed, the method comprising:

[0005] Align the target audio with its reference audio to obtain an audio alignment result, wherein the target audio contains packet loss frames with missing data.

[0006] Based on the packet loss frame position of the target audio and the audio alignment result, a first reference frame corresponding to the packet loss frame position in the reference audio is obtained, and a second reference frame located before and adjacent to the first reference frame is obtained.

[0007] Based on the audio feature association between the first reference frame and the second reference frame, data recovery is performed on the lost frame using the preceding target frame that is located before and adjacent to the lost frame as a reference.

[0008] According to one aspect of the embodiments of this application, an audio packet loss data recovery device is disclosed, the device comprising:

[0009] An audio alignment module is configured to align a target audio and a reference audio of the target audio to obtain an audio alignment result, wherein the target audio contains packet loss frames with missing data.

[0010] The reference frame acquisition module is configured to acquire a first reference frame in the reference audio corresponding to the packet loss frame position based on the packet loss frame position of the target audio and the audio alignment result, and acquire a second reference frame located before the first reference frame and adjacent to the first reference frame.

[0011] The data recovery module is configured to perform data recovery on the lost frame based on the audio feature association between the first reference frame and the second reference frame, using the preceding target frame located before and adjacent to the lost frame as a reference.

[0012] According to one aspect of the embodiments of this application, an electronic device is disclosed, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the methods provided in the various optional implementations described above.

[0013] According to one aspect of the embodiments of this application, a computer program medium is disclosed, on which computer-readable instructions are stored, which, when executed by a computer's processor, cause the computer to perform the methods provided in the various optional implementations described above.

[0014] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.

[0015] In this application, the target audio and reference audio with lost frames are aligned. Then, based on the lost frame position of the target audio and the audio alignment result, a first reference frame corresponding to the lost frame position in the reference audio is obtained, and a second reference frame preceding and adjacent to the first reference frame is also obtained. Then, based on the audio feature association between the first and second reference frames, and using the preceding target frame of the lost frame as a reference, data recovery of the lost frame is performed. Using this method, regardless of whether the signals of adjacent frames have changed significantly, this application can accurately recover data from lost frames, meeting the accuracy requirements of audio packet loss data recovery. Simultaneously, since this application saves the waiting time spent obtaining subsequent information after the lost frame, it also meets the low latency requirements of audio packet loss data recovery. Therefore, this application can simultaneously meet the accuracy and low latency requirements of audio packet loss data recovery.

[0016] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0017] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application. Attached Figure Description

[0018] The above and other objectives, features and advantages of this application will become more apparent from a detailed description of exemplary embodiments thereof with reference to the accompanying drawings.

[0019] Figure 1 A schematic diagram of an exemplary system architecture according to an embodiment of this application is shown.

[0020] Figure 2 A flowchart of an audio packet loss data recovery method according to an embodiment of this application is shown.

[0021] Figure 3 A schematic diagram of an audio alignment result according to an embodiment of this application is shown.

[0022] Figure 4 A schematic diagram of the target audio and reference audio before alignment according to an embodiment of this application is shown.

[0023] Figure 5 A schematic diagram of aligned target audio and reference audio according to an embodiment of this application is shown.

[0024] Figure 6 The diagram illustrates a process for implementing multi-person online chorus based on the audio packet loss data recovery method provided in this application, according to one embodiment of this application.

[0025] Figure 7 A block diagram of an audio packet loss data recovery apparatus according to an embodiment of this application is shown.

[0026] Figure 8 A hardware diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation

[0027] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided to make the description of this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The drawings are merely illustrative of this application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0028] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more exemplary embodiments. Numerous specific details are provided in the following description to give a full understanding of exemplary embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced with one or more of the specific details omitted, or other methods, components, steps, etc., can be employed. In other instances, well-known structures, methods, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0029] Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0030] This application provides an audio packet loss data recovery method, which mainly aims to recover data from lost frames in the target audio that have missing data, thereby ensuring the intelligibility and fluency of the target audio.

[0031] Furthermore, the audio packet loss data recovery method provided in this application is mainly for recovering audio packet loss that occurs during online network audio transmission. It meets the accuracy requirements of packet loss data recovery while also meeting the low latency requirements of packet loss data recovery.

[0032] It's important to note that during network transmission, packet loss is unavoidable. This can occur due to physical loss of data packets, or because network jitter and transmission delays cause packets to arrive late and be discarded due to exceeding the acceptable range for the actual business. Therefore, in online audio transmission, it's inevitable that some audio will contain missing data frames, resulting in blank segments in the transmitted audio. Understandably, audio with blank segments is more difficult for users to understand and is less smooth than complete audio. The more packet loss frames, the more choppy the audio becomes.

[0033] Therefore, to ensure the intelligibility and fluency of the transmitted audio, it is necessary to recover the lost audio data, i.e., the lost audio frames. The technique for recovering lost audio data is called packet loss recovery technology, or packet loss concealment technology (PLC), and it is mainly used to reconstruct the location signal of the lost frames.

[0034] Figure 1 A schematic diagram of an exemplary system architecture for a multi-person online chorus application according to an embodiment of this application is shown.

[0035] like Figure 1 As shown, the execution entity of the audio packet loss data recovery method in this system architecture can be the terminal of user 11 and user 12, or server 20. User 11's terminal includes one or more of a portable computer 111, a tablet computer 112, and a smartphone 113; user 12's terminal includes one or more of a portable computer 121, a tablet computer 122, and a smartphone 123. Server 20 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0036] In a multi-person online chorus application, User 11 and User 12 are singers who sing the same song together online. User 11 and User 12 are merely examples to demonstrate that multiple singers can sing the same song together, and do not imply that there can only be two singers singing the same song together.

[0037] In a multi-user online chorus application, the audio recordings of user 11 and user 12 need to be mixed to create the auditory effect of users 11 and 12 singing the same song together. Furthermore, during the mixing process, packet loss data recovery must be performed on the audio recordings to ensure intelligibility and smoothness.

[0038] In this method, server 20 can act as the execution entity for the audio packet loss data recovery. After receiving the audio voice of user 11, server 20 first performs audio packet loss data recovery, then aligns and mixes the audio, and then transmits the mixed audio to user 12's terminal. Similarly, after receiving the audio voice of user 12, server 20 first performs audio packet loss data recovery, then aligns and mixes the audio, and then transmits the mixed audio to user 11's terminal.

[0039] Alternatively, the user's terminal can act as the execution entity for the audio packet loss data recovery method. After receiving the audio of user 12's voice from server 20, user 11's terminal first performs audio packet loss data recovery, then aligns and mixes the audio, and finally outputs it locally. Similarly, after receiving the audio of user 11's voice from server 20, user 12's terminal first performs audio packet loss data recovery, then aligns and mixes the audio, and finally outputs it locally.

[0040] Understandably, multi-person online singing applications have high requirements for the accuracy of packet loss data recovery. If there is a large deviation in the recovery of packet loss data, it will lead to audio stuttering or distortion. In addition, multi-person online singing applications have very high requirements for latency. If the latency is slightly too high, users can easily perceive poor singing quality (for example, the two people singing the same content are out of sync, or the rhythm is out of sync when the two people sing together).

[0041] The audio packet loss data recovery method provided in this application is mainly proposed for multi-person online chorus applications, and simultaneously meets the accuracy and latency requirements of multi-person online chorus applications for packet loss data recovery.

[0042] It should be noted that although the audio packet loss data recovery method provided in this application is mainly proposed for multi-person online chorus applications, it is not limited to multi-person online chorus applications. Therefore, the system architecture shown in this embodiment should not limit the function and scope of use of this application.

[0043] Figure 2 A flowchart of an audio packet loss data recovery method provided in an embodiment of this application is shown. An exemplary execution entity of this method is a server, and the method includes:

[0044] Step S210: Align the target audio and the reference audio of the target audio to obtain the audio alignment result, wherein the target audio contains packet loss frames with missing data.

[0045] Step S220: Based on the packet loss frame position and audio alignment result of the target audio, obtain the first reference frame corresponding to the packet loss frame position in the reference audio, and obtain the second reference frame located before and adjacent to the first reference frame.

[0046] Step S230: Based on the audio feature association between the first reference frame and the second reference frame, and taking the preceding target frame that is located before and adjacent to the lost frame as a reference, perform data recovery on the lost frame.

[0047] Specifically, in this embodiment of the application, if the received audio is found to contain a lost frame with missing data, then the audio is used as the target audio for data recovery of the lost frame.

[0048] In related technologies, data recovery for lost frames is based on the frames preceding and following the lost frame in the target audio, using interpolation methods to recover the data. For example, let the lost frame in the target audio be si, where i is an integer representing the frame number. Then, related technologies recover data from the lost frame si by using interpolation methods based on the frame preceding si (i-1) and the frame following si.

[0049] The methods employed by related technologies are based on the premise that adjacent audio frames exhibit strong characteristic correlations; that is, if the audio is short-term stationary, then adjacent audio frames are also stationary. However, this premise is not always true—adjacent audio frames are not always stationary; significant differences can occur when switching between different phonemes; even within the same phoneme, adjacent frames can differ considerably. When adjacent frame signals change significantly, the premise upon which the methods rely weakens, causing the methods to fail to accurately capture changes between adjacent frames, thus reducing the accuracy of packet loss recovery. Furthermore, the more consecutive frames of packet loss occur, the lower the accuracy of packet loss recovery becomes. Typically, once the number of consecutive lost frames exceeds three, the related technologies become completely unable to recover any of the lost frames.

[0050] Furthermore, when the methods employed by the related technologies are applied to multi-person online chorus applications, these applications have very high latency requirements. Therefore, they do not allow sufficient time for the technologies to acquire subsequent frames as references for packet loss data recovery. Thus, to meet the low latency requirements of multi-person online chorus applications, the technologies can only use previous frames as references for packet loss data (e.g., recovering data from the lost frame si based solely on the frame s(i-1) preceding the lost frame si in the target audio). This further reduces the accuracy of packet loss data recovery. Conversely, to ensure accuracy in packet loss data recovery, the technologies cannot meet the latency requirements of multi-person online chorus applications. Therefore, it is evident that the methods employed by the related technologies cannot simultaneously satisfy both the accuracy and low latency requirements for packet loss data recovery.

[0051] To simultaneously meet the accuracy and low latency requirements of packet loss data recovery, this application aligns the target audio and its reference audio after receiving the target audio, obtaining an audio alignment result. The reference audio refers to the audio that serves as a reference template for the target audio. For example, in a multi-person online chorus, if the target audio is a recording of a user singing "Two Tigers," then the original vocals of "Two Tigers" can be used as the reference audio.

[0052] Audio alignment results are used to describe target frames and reference frames that are in the same position within their respective audio sequences. A target frame refers to an audio frame in the target audio sequence, and a lost frame specifically refers to an audio frame in the target audio sequence with missing data. A reference frame refers to an audio frame in the reference audio sequence.

[0053] After obtaining the audio alignment result, by combining it with the packet loss frame position of the target audio, the first reference frame corresponding to the packet loss frame position in the reference audio can be obtained, and the second reference frame located before and adjacent to the first reference frame can be obtained. Figure 3 The diagram shows the audio alignment result. The i-th frame of the target audio s is lost, i.e., the lost frame of the target audio is si. Aligning the segments containing the same information in the target audio s and the reference audio m yields the following result: Figure 3 The audio alignment result is shown. Figure 3 In the diagram, the two ends of the dashed line represent the target frame and the reference frame, which are in the same position in their respective audio segments. Therefore, after obtaining the audio alignment result, the first reference frame mi corresponding to the packet loss frame si in the reference audio m can be obtained based on the packet loss frame si of the target audio s, and the second reference frame m(i-1) can be obtained.

[0054] Continue to refer to Figure 3 The diagram illustrates the audio alignment results. Since the reference audio m is the reference template for the target audio s, the audio feature association between the lost frame si and its adjacent preceding target frame s(i-1) is highly similar to the audio feature association between the first reference frame mi and the second reference frame m(i-1). Therefore, regardless of whether the signals of adjacent frames change significantly, this application, based on the audio feature association between the first and second reference frames and using the preceding target frame corresponding to the lost frame as a reference, can accurately recover the data from the lost frame, meeting the accuracy requirements for audio packet loss data recovery.

[0055] Furthermore, since this application recovers data from lost frames based on prior knowledge no later than the lost frame (the first and second reference frames of the reference audio, and the preceding target frame of the lost frame), it is not necessary to obtain subsequent information after the lost frame (e.g., the following target frame of the lost frame). Therefore, this application saves the waiting time spent obtaining subsequent information after the lost frame, thereby also meeting the low latency requirement for audio lost frame data recovery.

[0056] Therefore, in this application, the target audio and reference audio with lost frames are aligned. Then, based on the lost frame position of the target audio and the audio alignment result, a first reference frame corresponding to the lost frame position in the reference audio is obtained, and a second reference frame preceding and adjacent to the first reference frame is obtained. Then, based on the audio feature association between the first and second reference frames, and using the preceding target frame of the lost frame as a reference, data recovery of the lost frame is performed. Using this method, regardless of whether the signals of adjacent frames have changed significantly, this application can accurately recover data from lost frames, meeting the accuracy requirements of audio packet loss data recovery. Simultaneously, since this application saves the waiting time spent obtaining subsequent information after the lost frame, it also meets the low latency requirements of audio packet loss data recovery. Therefore, this application can simultaneously meet the accuracy and low latency requirements of audio packet loss data recovery.

[0057] In one embodiment, the audio packet loss data recovery method provided in this application further includes:

[0058] The target audio is received using the User Datagram Protocol (UDP).

[0059] In this embodiment, data transmission between the server and the user's terminal is performed using the User Datagram Protocol (UDP). Specifically, after the terminal records the user's voice audio, it compresses and encodes it, and then sends the compressed and encoded data to the server via UDP. The server receives the compressed and encoded data via UDP, decodes it to obtain the user's voice audio, and then performs packet loss detection on the user's voice audio. If packet loss is detected in the user's voice audio, it is used as the target audio, and the method provided in this application is used to recover the packet loss data.

[0060] It should be noted that the reason UDP is used for audio data transmission in this embodiment is that UDP has the characteristic of low data transmission latency, which can save the latency caused by audio data transmission. Although UDP is not reliable and is prone to packet loss, the method provided in this application can quickly and accurately recover the lost target audio data, thus compensating for the shortcomings of UDP caused by packet loss. Overall, it ensures accuracy while further reducing latency.

[0061] In one embodiment, the audio packet loss data recovery method provided in this application further includes:

[0062] When the target audio is the user's vocal audio obtained by recording the user singing, the original vocal audio of the song sung by the user is obtained and used as the reference audio.

[0063] This embodiment primarily addresses online singing applications. It should be noted that the online singing applications in this embodiment include: single-person online singing applications (e.g., a single streamer singing a song online in a live broadcast room) and multi-person online duet applications (e.g., multiple singers singing a song together online in a karaoke client).

[0064] In this embodiment, the terminal of the online singing application records the user's voice audio, compresses and encodes it, and then sends the compressed and encoded data to the server. After the server decodes the user's voice audio, it performs packet loss detection on the user's voice audio. If the server detects packet loss in the user's voice audio, it uses that audio as the target audio and uses the original vocal audio of the song sung by the user as the reference audio.

[0065] In one embodiment, the audio packet loss data recovery method provided in this application further includes:

[0066] When the target audio is the first user's voice audio obtained by recording the first user singing, the second user's voice audio obtained by recording the second user singing is acquired, and the second user's voice audio is used as the reference audio. The content of the first user and the second user singing synchronously includes the packet loss frame position.

[0067] This embodiment primarily addresses multi-user online chorus applications. It should be noted that, in this embodiment, "chorus" in multi-user online chorus applications refers to a broad, song-based performance method. Specifically, in this embodiment, chorus in multi-user online chorus applications can be divided into two categories: different users singing different segments of the same song, which can be called duet; and different users simultaneously singing the same segment of the same song, which can be called unison singing.

[0068] In this embodiment, for a synchronized singing segment, the audio of other users' voices can be used as reference audio. This is because, for a synchronized singing segment, the server receives the audio of multiple users singing that segment, and the audio characteristics of these users' voices are highly similar. Even if one user's voice audio is lost, the probability of another user's voice audio being lost at the same frame position is very low. Therefore, the other user's voice audio can be used as a reference for recovering lost data.

[0069] Specifically, the server receives the audio recording of the first user's singing from the first user's terminal. If packet loss detection confirms that there is packet loss in the first user's audio recording, and the lost frame is located in the same singing segment, then the server identifies the second user who sang the same singing segment with the first user, obtains the audio recording of the second user's singing from the second user's terminal, and uses the second user's audio recording as a reference audio.

[0070] For example, in a song sung by Xiaoming and Xiaohong, the segment from 1 minute 30 seconds to 1 minute 50 seconds requires them to sing in unison. After receiving Xiaoming's audio, the server confirms that there is packet loss at 1 minute 40 seconds. Since 1 minute 40 seconds falls within the segment where they should be singing in unison, the server can use Xiaohong's audio as a reference to recover the packet loss data from Xiaoming's audio.

[0071] It should be noted that for choral singing segments in multi-user online chorus applications, the original singer's vocals can also be used as reference audio. The reason this embodiment uses other users' vocals as reference audio is primarily because, in some cases, the server may not be able to obtain the original singer's vocals (e.g., the user-sung song is an original song not publicly available, and there is no publicly available original vocal audio). To enable the server to perform packet loss data recovery according to the method provided in this application even without the original vocal audio, this embodiment proposes using other users' vocals as reference audio, thereby expanding the applicability of packet loss data recovery.

[0072] In one embodiment, aligning the target audio and a reference audio of the target audio to obtain an audio alignment result includes:

[0073] Obtain the target audio fingerprint of the target audio and the reference audio fingerprint of the reference audio.

[0074] Based on the target audio fingerprint and the reference audio fingerprint, the target audio and the reference audio are aligned to obtain the audio alignment result.

[0075] In this embodiment, an audio fingerprint matching method is used to align the target audio and the reference audio.

[0076] Specifically, the frequency domain power spectrum of each target frame in the target audio can be calculated, and then the target audio fingerprint can be calculated based on the frequency domain power spectrum of each target frame; similarly, the frequency domain power spectrum of each reference frame in the reference audio can be calculated, and then the reference audio fingerprint can be calculated based on the frequency domain power spectrum of each reference frame.

[0077] After calculating the target and reference audio fingerprints, the distance between them can be measured by comparing their distances (e.g., Euclidean distance). The smaller the distance, the closer the target and reference audio fingerprints are. The key to aligning them lies in finding the position that makes them closest, and then using that position as a reference to associate the target frame with the reference frame.

[0078] To find the closest position between the two, the reference audio can be offset frame by frame, and the reference audio fingerprints corresponding to multiple consecutive frames of the offset reference audio can be extracted. Then, the distance between the reference audio fingerprints corresponding to multiple consecutive frames of the offset reference audio and the target audio fingerprints corresponding to multiple consecutive frames of the target audio can be calculated. Since this distance is calculated based on the reference audio fingerprints corresponding to multiple consecutive frames of the offset reference audio, this distance is described as the offset distance.

[0079] By performing a series of operations, including offsetting the audio, extracting the offset audio fingerprint, and calculating the offset distance, the position that minimizes the offset distance is selected. This position is the one that makes the target audio fingerprint and the reference audio fingerprint the closest, thus achieving alignment between the target audio and the reference audio.

[0080] In one embodiment, aligning the target audio and a reference audio of the target audio to obtain an audio alignment result includes:

[0081] Obtain the target audio melody of the target audio and the reference audio melody of the reference audio.

[0082] Based on the target audio melody and the reference audio melody, the target audio and the reference audio are aligned to obtain the audio alignment result.

[0083] In this embodiment, a humming recognition method is used to align the target audio and the reference audio.

[0084] Specifically, MIDI (Musical Instrument Digital Interface) extraction technology can be used to extract the target audio melody and the reference audio melody of the reference audio.

[0085] Since both the target audio melody and the reference audio melody are time series data, DTW (Dynamic Time Warping) technology can be used to locally scale the target audio melody and the reference audio melody on the time axis to calculate the similarity between the target audio melody and the reference audio melody, and then find the position that makes the target audio and the reference audio the closest, thereby aligning the target audio and the reference audio.

[0086] Figures 4 to 5 This diagram illustrates an embodiment of aligning target audio and reference audio according to this application. Specifically, Figure 4 A schematic diagram of the target audio and reference audio before alignment is shown according to an embodiment of this application. Figure 5 An embodiment of this application is shown. Figure 4 A schematic diagram of the aligned target audio and reference audio in the embodiment.

[0087] See Figures 4 to 5 In one embodiment, the server obtains as follows: Figure 4 The target audio s and reference audio m are shown. Considering that audio alignment is mainly achieved by matching existing target and reference frames, lost frames in the target audio s have virtually no direct impact on the audio alignment process. Figure 4 The packet loss frames in the target audio s are not shown.

[0088] By employing audio fingerprint matching or humming recognition to align the target audio s and the reference audio m, it can be confirmed that the k-th frame of the target audio s and the j-th frame of the reference audio m are most similar. Therefore, by aligning the target frame sk of the target audio s with the reference frame mj of the reference audio m, alignment between the target audio s and the reference audio m is achieved, resulting in the following... Figure 5 The aligned target audio s and reference audio m are shown.

[0089] In one embodiment, based on the audio feature association between the first reference frame and the second reference frame, and using the preceding target frame that is located before and adjacent to the lost frame as a reference, data recovery of the lost frame is performed, including:

[0090] Obtain the ratio between the audio features of the first reference frame and the audio features of the second reference frame;

[0091] The audio features of the lost frame are calculated based on the ratio and the product of the audio features of the previous target frame.

[0092] Data recovery is performed on the lost frames based on their audio characteristics.

[0093] In this embodiment, the audio features used for data recovery of lost frames include, but are not limited to: line spectrum pairs (lsp), pitch period (pitch), gain, etc.

[0094] Taking the pitch period as an example, let the pitch period of the first reference frame mi be p_mi, the pitch period of the second reference frame m(i-1) be p_m(i-1), the pitch period of the lost frame si be p_si, and the pitch period of the preceding target frame s(i-1) of the lost frame be p_s(i-1). The pitch period p_si of the lost frame si is a value to be determined, and can be calculated using the following formula.

[0095] p_si = p_s(i-1) * p_mi / p_m(i-1)

[0096] Similarly, the pitch period can be calculated to obtain the audio features such as the line spectrum pairs and gain of the lost frame. Then, based on the calculated pitch period, line spectrum pairs, gain and other audio features, the audio signal of the lost frame can be decoded and recovered.

[0097] It should be noted that the audio feature association between the first reference frame and the second reference frame is actually characterized based on the functional relationship between the audio features, and therefore is not limited to being represented as the ratio between audio features.

[0098] Figure 6 This illustration shows a flowchart of a multi-person online chorus based on the audio packet loss data recovery method provided in this application, according to an embodiment of this application.

[0099] See Figure 6 In this embodiment, the client on the user's terminal records the user's singing to obtain the user's voice audio, compresses and encodes it, and then transmits the compressed and encoded data to the mixing server via UDP over the network.

[0100] The mixing server receives the compressed and encoded data, decodes it to obtain the user's vocal audio. Then, it searches and aligns the user's vocal audio with the original vocal audio of the song sung by the user to obtain the audio alignment result.

[0101] Furthermore, after the mixing server receives the user's audio, it performs packet loss detection to confirm whether there is any packet loss in the user's audio.

[0102] If there is no packet loss in the user's voice audio, then based on the audio alignment result, the user's voice audio is mixed with the accompaniment audio of the song sung by the user, or the user's voice audio is mixed with the voice audio of other users to achieve a chorus effect.

[0103] If packet loss occurs in the user's audio, a packet loss hiding algorithm based on prior knowledge, i.e., based on the audio packet loss data recovery method provided in this application, is used to recover the packet loss data from the user's audio. After the packet loss data recovery is completed, the user's audio is mixed with the accompaniment audio of the song sung by the user, or the user's audio is mixed with the audio of other users, to achieve a chorus effect, based on the audio alignment result. The prior knowledge includes, but is not limited to: the first reference frame in the original audio corresponding to the packet loss frame position, and the second reference frame located before and adjacent to the first reference frame, and the preceding target frame in the user's audio located before and adjacent to the packet loss frame.

[0104] Figure 7 A block diagram of an audio packet loss data recovery apparatus according to an embodiment of this application is shown, the apparatus comprising:

[0105] The audio alignment module 310 is configured to align the target audio and the reference audio of the target audio to obtain an audio alignment result, wherein the target audio has packet loss frames with missing data.

[0106] The reference frame acquisition module 320 is configured to acquire a first reference frame in the reference audio corresponding to the packet loss frame position based on the packet loss frame position of the target audio and the audio alignment result, and acquire a second reference frame located before the first reference frame and adjacent to the first reference frame.

[0107] The data recovery module 330 is configured to perform data recovery on the lost frame based on the audio feature association between the first reference frame and the second reference frame, using the preceding target frame located before and adjacent to the lost frame as a reference.

[0108] In one exemplary embodiment of this application, the device is configured as follows:

[0109] When the target audio is a user's voice audio obtained by recording the user singing, the original vocal audio of the song sung by the user is obtained and the original vocal audio is used as the reference audio.

[0110] In one exemplary embodiment of this application, the device is configured as follows:

[0111] When the target audio is the first user's voice audio obtained by recording the first user singing, the second user's voice audio obtained by recording the second user singing is acquired, and the second user's voice audio is used as the reference audio, wherein the content of the first user and the second user singing synchronously includes the packet loss frame position.

[0112] In one exemplary embodiment of this application, the audio alignment module is configured as follows:

[0113] Obtain the target audio fingerprint of the target audio, and obtain the reference audio fingerprint of the reference audio;

[0114] Based on the target audio fingerprint and the reference audio fingerprint, the target audio and the reference audio are aligned to obtain the audio alignment result.

[0115] In one exemplary embodiment of this application, the audio alignment module is configured as follows:

[0116] Obtain the target audio melody of the target audio, and obtain the reference audio melody of the reference audio;

[0117] Based on the target audio melody and the reference audio melody, the target audio and the reference audio are aligned to obtain the audio alignment result.

[0118] In an exemplary embodiment of this application, the data recovery module is configured as follows:

[0119] Obtain the ratio between the audio features of the first reference frame and the audio features of the second reference frame;

[0120] The audio features of the lost frame are calculated based on the product of the ratio and the audio features of the previous target frame.

[0121] Data recovery is performed on the lost packet frames based on their audio characteristics.

[0122] In one exemplary embodiment of this application, the device is configured as follows:

[0123] The target audio is received using the User Datagram Protocol (UDP).

[0124] The following is for reference. Figure 8 To describe the electronic device 40 according to an embodiment of this application. Figure 8 The electronic device 40 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0125] like Figure 8 As shown, the electronic device 40 is manifested in the form of a general-purpose computing device. The components of the electronic device 40 may include, but are not limited to: at least one processing unit 410, at least one storage unit 420, and a bus 430 connecting different system components (including storage unit 420 and processing unit 410).

[0126] The storage unit stores program code that can be executed by the processing unit 410, causing the processing unit 410 to perform the steps described in the exemplary method description section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 410 can perform actions such as... Figure 2 The steps shown are as follows.

[0127] Storage unit 420 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 4201 and / or cache memory 4202, and may further include a read-only memory (ROM) 4203.

[0128] Storage unit 420 may also include a program / utility 4204 having a set (at least one) program module 4205, such program module 4205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0129] Bus 430 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0130] Electronic device 40 can also communicate with one or more external devices 500 (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 40, and / or any device that enables electronic device 40 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 450. Input / output (I / O) interface 450 is connected to display unit 440. Furthermore, electronic device 40 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 460. As shown, network adapter 460 communicates with other modules of electronic device 40 via bus 430. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 40, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0131] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the method according to the embodiments of this application.

[0132] In an exemplary embodiment of this application, a computer-readable storage medium is also provided, on which computer-readable instructions are stored, which, when executed by a computer's processor, cause the computer to perform the methods described in the above method embodiments.

[0133] According to one embodiment of this application, a program product for implementing the methods in the above-described method embodiments is also provided. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this invention is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0134] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0135] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0136] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0137] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as JAVA and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0138] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0139] Furthermore, although the steps of the method in this application are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0140] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0141] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the appended claims.

Claims

1. A method for recovering audio packet loss data, characterized in that, The method includes: Align the target audio and the reference audio of the target audio to obtain the audio alignment result. The target audio has packet loss frames with missing data. When the target audio is the user's voice audio obtained by recording the user singing, the original vocal audio of the song sung by the user is obtained and the original vocal audio is used as the reference audio. Based on the packet loss frame position of the target audio and the audio alignment result, a first reference frame corresponding to the packet loss frame position in the reference audio is obtained, and a second reference frame located before and adjacent to the first reference frame is obtained. Based on the audio feature association between the first reference frame and the second reference frame, data recovery is performed on the lost frame using the preceding target frame that is located before and adjacent to the lost frame as a reference.

2. The method according to claim 1, characterized in that, The method further includes: When the target audio is the first user's voice audio obtained by recording the first user singing, the second user's voice audio obtained by recording the second user singing is acquired, and the second user's voice audio is used as the reference audio, wherein the content of the first user and the second user singing synchronously includes the packet loss frame position.

3. The method according to claim 1, characterized in that, Align the target audio and its reference audio to obtain the audio alignment result, including: Obtain the target audio fingerprint of the target audio, and obtain the reference audio fingerprint of the reference audio; Based on the target audio fingerprint and the reference audio fingerprint, the target audio and the reference audio are aligned to obtain the audio alignment result.

4. The method according to claim 1, characterized in that, Align the target audio and its reference audio to obtain the audio alignment result, including: Obtain the target audio melody of the target audio, and obtain the reference audio melody of the reference audio; Based on the target audio melody and the reference audio melody, the target audio and the reference audio are aligned to obtain the audio alignment result.

5. The method according to claim 1, characterized in that, Based on the audio feature association between the first reference frame and the second reference frame, and using the preceding target frame that is located before and adjacent to the lost packet frame as a reference, data recovery is performed on the lost packet frame, including: Obtain the ratio between the audio features of the first reference frame and the audio features of the second reference frame; The audio features of the lost frame are calculated based on the product of the ratio and the audio features of the previous target frame. Data recovery is performed on the lost packet frames based on their audio characteristics.

6. The method according to claim 1, characterized in that, The method further includes: The target audio is received using the User Datagram Protocol (UDP).

7. An audio packet loss data recovery device, characterized in that, The device includes: An audio alignment module is configured to align a target audio and a reference audio of the target audio to obtain an audio alignment result. The target audio contains packet loss frames with missing data. When the target audio is a user's voice audio obtained by recording a user singing, the original vocal audio of the song sung by the user is obtained and the original vocal audio is used as the reference audio. The reference frame acquisition module is configured to acquire a first reference frame in the reference audio corresponding to the packet loss frame position based on the packet loss frame position of the target audio and the audio alignment result, and acquire a second reference frame located before the first reference frame and adjacent to the first reference frame. The data recovery module is configured to perform data recovery on the lost frame based on the audio feature association between the first reference frame and the second reference frame, using the preceding target frame located before and adjacent to the lost frame as a reference.

8. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, It stores computer-readable instructions that, when executed by the processor of a computer, cause the computer to perform the method described in any one of claims 1 to 6.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method as described in any one of claims 1 to 6.