Audio data generation method and apparatus, device, and storage medium

By acquiring historical device latency and playback device type, calculating and fusing candidate device latency data to align background music and voice data, the problem of inaccurate audio data processing caused by device latency differences in virtual space is solved, achieving higher precision audio data generation and a better virtual space communication experience.

CN116612772BActive Publication Date: 2026-04-21BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-12
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In virtual space scenarios, differences in device latency make it difficult to align background music and human voices. In existing technologies, inaccurate latency settings result in low accuracy in audio data processing.

Method used

By acquiring historical device latency data and playback device type, candidate device latency data is calculated and fused with historical device latency data to determine target device latency data to align with background music and voice data.

Benefits of technology

It improves the accuracy of device latency determination, makes background music and voice data more accurately aligned, enhances the accuracy of audio data generation and virtual space communication effects, and reduces audio data processing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612772B_ABST
    Figure CN116612772B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio data generation method and device, electronic equipment and storage medium. The method comprises: in the case of determining that the account of the client enters the target virtual space, obtaining historical device delay data and the device type of the playing device currently playing background music of the client; the historical device delay data is the device delay data corresponding to the last time the account enters the target virtual space; according to the device type of the playing device, the delay between the background music and the voice data generated by the account entering the target virtual space is determined, and the candidate device delay data of the client in the target virtual space is obtained; the candidate device delay data and the historical device delay data are fused to obtain the target device delay data of the client in the target virtual space; and the target audio data is obtained by aligning the background music and the voice data according to the target device delay data. The present disclosure can improve the generation accuracy of the target audio data and reduce the processing cost of the audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio signal processing technology, and in particular to an audio data generation method, apparatus, device, and storage medium. Background Technology

[0002] In virtual spaces (e.g., RTC karaoke rooms), there is a delay between the background music (BGM) being played from the speakers and being picked up by the microphone. Device latency often severely affects the singing experience, causing listeners to hear the background music accompaniment faster than the vocals. Because different mobile devices have different hardware, their latency also varies. Ensuring the alignment of background music and vocal data across different types of mobile phones is a current challenge and pain point for RTC karaoke room audio technology.

[0003] In related technologies, solutions for aligning background music and vocals typically involve inserting a corresponding length of empty data at the beginning of the mixing stage, referencing a set delay value. This offsets the overall background music data, ensuring accurate mixing and alignment between the background music and vocal data. However, setting the delay value too small results in the background music data running faster than the vocal data, while setting it too large results in the vocal data running faster than the background music data, thus reducing the accuracy of audio data processing. Summary of the Invention

[0004] This disclosure provides an audio data generation method, apparatus, device, and storage medium to at least solve the problem in related technologies where setting the delay value too small causes background music data to outpace human voice data, while setting it too large causes human voice data to outpace background music data, thereby reducing the accuracy of audio data processing. The technical solution of this disclosure is as follows:

[0005] According to a first aspect of the present disclosure, an audio data generation method is provided, comprising:

[0006] Once it is confirmed that the client's account has entered the target virtual space, historical device latency data and the device type of the playback device are obtained; the historical device latency data is the device latency data corresponding to the last time the account entered the target virtual space; the playback device is the playback device of the background music currently being played by the client.

[0007] Based on the device type of the playback device, the delay between the background music and the voice data generated when the account enters the target virtual space is determined, and the candidate device delay data of the client in the target virtual space is obtained;

[0008] By fusing the candidate device latency data and the historical device latency data, the target device latency data of the client in the target virtual space is obtained;

[0009] The background music and the voice data are aligned based on the target device delay data to obtain the target audio data.

[0010] In an optional embodiment, fusing the candidate device latency data and the historical device latency data to obtain the target device latency data of the client in the target virtual space includes:

[0011] Determine the weight information corresponding to the candidate device delay data and the historical device delay data respectively;

[0012] Based on the weight information, the candidate device delay data and the historical device delay data are weighted and aggregated to obtain the target device delay data.

[0013] In an optional embodiment, the method further includes:

[0014] Upon determining that the account is entering the target virtual space for the first time, obtain the client's operating system information and the initial device latency data corresponding to the client's operating system information;

[0015] Based on the initial device delay data, the background music and the voice data generated when the account first enters the target virtual space are aligned to obtain the target audio data.

[0016] In an optional embodiment, the method for generating the historical device delay data includes:

[0017] When the account enters the target virtual space for the second time, the initial device latency data is determined to be historical device latency data, and the candidate device latency data and historical device latency data of the client in the target virtual space are merged to obtain the target device latency data;

[0018] The historical device delay data is updated based on the target device delay data, and the updated device delay data is re-determined as the historical device delay data.

[0019] When the account enters the target virtual space again, the process of merging the candidate device latency data and historical device latency data of the client in the target virtual space is repeated until the updated device latency data is re-determined as historical device latency data, until the account exits the target virtual space.

[0020] In an optional embodiment, determining the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtaining the candidate device delay data of the client in the target virtual space, includes:

[0021] When the account's time in the target virtual space meets a preset time condition, the time is processed in segments to obtain at least two time periods;

[0022] The candidate device latency data is obtained based on the latency between the background music and the voice data of the account in each time period, as well as the historical average device latency data corresponding to each time period.

[0023] The historical average device latency data corresponding to each time period represents the average latency information of the client's device latency data within the historical time period, and the historical time period is the time period preceding each of the at least two time periods; the latency between the background music and the account's voice data within each time period is determined based on the device type of the playback device.

[0024] In an optional embodiment, obtaining the candidate device latency data based on the latency between the background music and the account's voice data in each time period, and the historical average device latency data corresponding to each time period, includes:

[0025] Sort the at least two time periods in ascending order to obtain a time period sequence;

[0026] Based on the device type of the playback device, determine the delay between the background music and the voice data of the account in the time period that is ranked first in the time period sequence, and obtain the current historical average device delay data;

[0027] The remaining time period that is ranked first in the remaining time period sequence is determined as the current time period, and the current time period is deleted from the remaining time period sequence; the remaining time period sequence is the sequence of time periods excluding the time period ranked first.

[0028] Based on the device type of the playback device, determine the delay between the background music and the voice data of the account in the current time period, and obtain the current device delay data of the client in the current time period;

[0029] The current historical average device latency data and the current device latency data are merged, and the fusion result is re-determined as the current historical average device latency data;

[0030] Repeat the process of determining the remaining time period that is first in the remaining time period sequence as the current time period, until the fusion result is re-determined as the current historical average device latency data, until the current device latency data of the client in the remaining time period at the end of the sequence is fused to obtain the candidate device latency data.

[0031] In an optional embodiment, the current time period includes a first time period and a second time period, wherein the playback device for the client to play background music during the first time period is a speaker, and the playback device for the client to play background music during the second time period is headphones;

[0032] The step of determining the delay between the background music and the account's voice data within the current time period based on the device type of the playback device, and obtaining the current device delay data of the client within the current time period, includes:

[0033] Based on the echo cancellation delay determination information corresponding to the speaker, the delay between the background music and the voice data of the account in the first time period is determined, and the first device delay data of the client in the first time period is obtained; based on the base frequency delay determination information corresponding to the headphones, the delay between the background music and the voice data of the account in the second time period is determined, and the second device delay data of the client in the second time period is obtained.

[0034] Based on the time ratio information of the first time period and the second time period to the current time period, the delay data of the first device and the delay data of the second device are fused to obtain the delay data of the current device.

[0035] In an optional embodiment, determining the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtaining the candidate device delay data of the client in the target virtual space, includes:

[0036] When the playback device is a headphone, the music digital interface data of the background music is acquired; the music digital interface data includes a first pitch information sequence of the background music.

[0037] The fundamental frequency of the voice data of the account within the target virtual space is detected to obtain the fundamental frequency detection result.

[0038] The fundamental frequency detection result is subjected to pitch conversion processing to obtain a second pitch information sequence;

[0039] The first pitch information in the first pitch information sequence is normalized to a preset scale to obtain a first target pitch information sequence. The second pitch information sequence in the second pitch information sequence is normalized to the preset scale to obtain a second target pitch information sequence.

[0040] The offset information is determined when the similarity between the first target pitch information sequence and the second target pitch information sequence meets a preset condition, and the candidate device delay data is obtained.

[0041] In an optional embodiment, determining the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtaining the candidate device delay data of the client in the target virtual space, includes:

[0042] When the playback device is a speaker, reference data for the background music and external audio data are acquired; the external audio data is the sound data collected from the external environment, including the echo data generated by the background music played inside the client after being diffused through the speaker and the voice data of the account in the target virtual space; the reference data is the original data of the background music played inside the client.

[0043] The reference data and the external audio data are input into a linear filter, and the echo data in the external audio data is filtered by the linear filter to obtain the filtered external audio data; the delay of the reference data and the filtered external audio data is determined to obtain the candidate device delay data.

[0044] According to a second aspect of the present disclosure, an audio data generation apparatus is provided, comprising:

[0045] The data acquisition module is configured to acquire historical device latency data and the device type of the playback device when it is determined that the client's account has entered the target virtual space; the historical device latency data is the device latency data corresponding to the last time the account entered the target virtual space; the playback device is the playback device of the background music currently being played by the client.

[0046] The delay calculation module is configured to determine the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtain the candidate device delay data of the client in the target virtual space;

[0047] The fusion module is configured to fuse the candidate device latency data and the historical device latency data to obtain the target device latency data of the client in the target virtual space;

[0048] The alignment module is configured to perform alignment of the background music and the voice data based on the target device delay data to obtain target audio data.

[0049] In an optional embodiment, the fusion module includes:

[0050] The weight information determination unit is configured to determine the weight information corresponding to the candidate device delay data and the historical device delay data respectively;

[0051] The weight aggregation determination unit is configured to perform weight aggregation processing on the candidate device delay data and the historical device delay data according to the weight information to obtain the target device delay data.

[0052] In an optional embodiment, the apparatus further includes:

[0053] The initial device delay data acquisition module is configured to acquire the client's operating system and the initial device delay data corresponding to the client's operating system when it is determined that the account is entering the target virtual space for the first time.

[0054] The target audio data generation module is configured to perform actions based on the initial device delay data, aligning the background music and the voice data generated when the account first enters the target virtual space, to obtain target audio data.

[0055] In an optional embodiment, the device further includes a historical device delay data generation module, which includes:

[0056] The target device latency data generation unit is configured to, when the account enters the target virtual space for the second time, determine that the initial device latency data is historical device latency data, and merge the candidate device latency data of the client currently in the target virtual space with the historical device latency data to obtain the target device latency data;

[0057] The re-determination unit is configured to perform an update of the historical device delay data based on the target device delay data, and re-determine the updated device delay data as the historical device delay data;

[0058] The repeat fusion unit is configured to perform the following operation when the account enters the target virtual space again: repeat the fusion of the client's current candidate device latency data and historical device latency data in the target virtual space until the updated device latency data is re-determined as historical device latency data, until the account exits the target virtual space.

[0059] In an optional embodiment, the delay calculation module includes:

[0060] The segmentation unit is configured to process the time in segments to obtain at least two time periods when the time of the account in the target virtual space meets a preset time condition.

[0061] The candidate device delay data generation unit is configured to perform the following: based on the delay between the background music and the voice data of the account in each time period, and the historical average device delay data corresponding to each time period, the candidate device delay data is obtained.

[0062] The historical average device latency data corresponding to each time period represents the average latency information of the client's device latency data within the historical time period, and the historical time period is the time period preceding each of the at least two time periods; the latency between the background music and the account's voice data within each time period is determined based on the device type of the playback device.

[0063] In an optional embodiment, the candidate device delay data generation unit includes:

[0064] An ascending subunit is configured to perform ascending sorting of the at least two time periods to obtain a time period sequence;

[0065] The current historical average device latency data generation subunit is configured to perform the following: determine the latency between the background music and the voice data of the account in the first time period of the time sequence according to the device type of the playback device, and obtain the current historical average device latency data.

[0066] The current time period determination subunit is configured to determine the first remaining time period in the remaining time period sequence as the current time period, and then delete the current time period from the remaining time period sequence; the remaining time period sequence is the sequence of time periods excluding the first remaining time period.

[0067] The current device delay data generation subunit is configured to determine the delay between the background music and the account's voice data in the current time period based on the device type of the playback device, and obtain the current device delay data of the client in the current time period;

[0068] The redefined subunit is configured to perform the fusion of the current historical average device latency data and the current device latency data, and redefine the fusion result as the current historical average device latency data;

[0069] The repeat execution subunit is configured to repeatedly perform the operation of determining the remaining time period at the first position of the remaining time period sequence as the current time period, until the fusion result is re-determined as the current historical average device latency data, until the current device latency data of the client in the remaining time period at the last position of the sequence is fused to obtain the candidate device latency data.

[0070] In an optional embodiment, the current time period includes a first time period and a second time period, wherein the playback device for the client to play background music during the first time period is a speaker, and the playback device for the client to play background music during the second time period is headphones;

[0071] This delay calculation module includes:

[0072] The device delay data generation unit is configured to perform the following actions: determine the delay between the background music and the voice data of the account in the first time period based on the echo cancellation delay determination information corresponding to the speaker, and obtain the first device delay data of the client in the first time period; determine the delay between the background music and the voice data of the account in the second time period based on the base frequency delay determination information corresponding to the headphones, and obtain the second device delay data of the client in the second time period.

[0073] The delay data fusion unit is configured to perform the following operations: fused the delay data of the first device and the delay data of the second device based on the time ratio information of the first time period and the second time period to the current time period, respectively, to obtain the delay data of the current device.

[0074] In an optional embodiment, the delay calculation module includes:

[0075] The music digital interface acquisition unit is configured to acquire music digital interface data of background music; the music digital interface data includes a first pitch information sequence of the background music.

[0076] The baseband detection unit is configured to perform baseband detection on the voice data of the account in the target virtual space, and obtain the baseband detection result;

[0077] A pitch conversion unit is configured to perform pitch conversion processing on the fundamental frequency detection result to obtain a second pitch information sequence;

[0078] The normalization unit is configured to perform normalization processing on the first pitch information in the first pitch information sequence to a preset scale to obtain a first target pitch information sequence, and to perform normalization processing on the second pitch information sequence in the second pitch information sequence to the preset scale to obtain a second target pitch information sequence.

[0079] The offset information determination unit is configured to calculate the similarity between the first target pitch information sequence and the second target pitch information sequence, and to obtain the offset information when a preset condition is met, thereby obtaining the candidate device delay data.

[0080] In an optional embodiment, the delay calculation module includes:

[0081] The reference audio data acquisition unit is configured to acquire reference data and external audio data of the background music; the external audio data is sound data collected from the external environment, including echo data generated by the background music played inside the client after being diffused through the speaker and voice data of the account in the target virtual space; the reference data is the original data of the background music played inside the client.

[0082] The filtering unit is configured to input the reference data and the external audio data to a linear filter, filter the echo data in the external audio data through the linear filter to obtain filtered external audio data, and determine the delay of the reference data and the filtered external audio data to obtain the candidate device delay data.

[0083] According to a third aspect of the present disclosure, an electronic device for generating audio data is provided, comprising:

[0084] processor;

[0085] Memory used to store the processor's executable instructions;

[0086] The processor is configured to execute the instructions to implement the audio data generation method as described in any of the above embodiments.

[0087] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device performs an audio data generation method as described in any of the above embodiments.

[0088] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the audio data generation method described in any of the above embodiments.

[0089] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0090] The audio data generation method, apparatus, device, and storage medium provided in the embodiments of this disclosure, when it is determined that a client's account has entered a target virtual space, acquires historical device latency data and the device type of the playback device; the historical device latency data is the device latency data corresponding to the last time the account entered the target virtual space; the playback device is the playback device currently playing background music on the client. Based on the device type of the playback device, the latency between the background music and the voice data within the target virtual space is determined, obtaining candidate device latency data for the client in the target virtual space; the candidate device latency data and the historical device latency data are fused to obtain target device latency data for the client in the target virtual space; the background music and voice data are aligned according to the target device latency data to obtain target audio data. As can be seen, in this embodiment, when the client enters the target virtual space, the device type of the playback device is used to accurately calculate the latency data of a candidate device. By fusing this latency data with historical device latency data, the client's device latency can be made closer to the client's actual device latency, improving the accuracy of the client's device latency determination. This allows for more precise alignment of background music and voice data, improving the generation accuracy of the target audio data and the real-time communication effect of the target virtual space (e.g., the singing effect in an RTC karaoke room), thereby enhancing the client's account experience. Furthermore, since the historical device latency data in this embodiment corresponds to the device type of the playback device, this embodiment can be applied to various manufacturers and types of client devices. Clients do not need to purchase various models of devices, reducing the cost of audio data processing and improving the efficiency of audio data processing.

[0091] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0092] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0093] Figure 1 This is an application environment diagram illustrating an audio data generation method according to an exemplary embodiment.

[0094] Figure 2This is a flowchart illustrating an audio data generation method according to an exemplary embodiment.

[0095] Figure 3 This is a schematic diagram illustrating the principle of echo cancellation delay determination information according to an exemplary embodiment.

[0096] Figure 4 This is a flowchart illustrating, according to an exemplary embodiment, a method for determining candidate device latency data for a client in a target virtual space.

[0097] Figure 5 This is a flowchart illustrating, according to an exemplary embodiment, a method for determining candidate device latency data for a client in a target virtual space.

[0098] Figure 6 This is a flowchart illustrating a method for determining the current device delay data of a client within the current time period, according to an exemplary embodiment.

[0099] Figure 7 This is a flowchart illustrating a method for generating historical device delay data according to an exemplary embodiment.

[0100] Figure 8 This is a flowchart illustrating a method for calculating delay data of a target device according to an exemplary embodiment.

[0101] Figure 9 This is a flowchart illustrating another method for generating audio data according to an exemplary embodiment.

[0102] Figure 10 This is a block diagram of an audio data generation apparatus according to an exemplary embodiment.

[0103] Figure 11 This is a block diagram illustrating an electronic device for generating audio data according to an exemplary embodiment. Detailed Implementation

[0104] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0105] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0106] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0107] The audio data generation method provided in this disclosure can be executed by an electronic device in an instant messaging system for processing audio data. Exemplarily, the electronic device can be at least one of at least two clients performing instant messaging. Exemplarily, the client can include a smartphone, desktop computer, tablet computer, laptop computer, wearable smart terminal, etc.

[0108] Please see Figure 1 , Figure 1 This is an application environment diagram illustrating an audio data generation method according to an exemplary embodiment. The application environment may include client 01 and client 02. Client 01 and client 02 can be at least two clients engaged in real-time communication. Client 02 can communicate with client 01 via wired or wireless means.

[0109] It should be noted that, Figure 1 This is merely one application environment for the audio data generation method disclosed herein; in practical applications, other application environments may also be included.

[0110] Figure 2 This is a flowchart illustrating an audio data generation method according to an exemplary embodiment, such as... Figure 2 As shown, the method may include the following steps.

[0111] In step S11, if it is determined that the client's account has entered the target virtual space, historical device latency data and the device type of the playback device are obtained; the historical device latency data is the device latency data corresponding to the last time the account entered the target virtual space; the playback device is the playback device of the background music currently being played by the client.

[0112] Optionally, the target virtual space can refer to a space where the client's account communicates instantly with other accounts. More specifically, the target virtual space can be a space where the client's account engages in real-time voice communication with other accounts. For example, the target virtual space could be an online karaoke room (e.g., an RTC karaoke room), a live streaming room, etc.

[0113] Optionally, in step S11 above, the situation where the client's account enters the target virtual space may refer to a situation where the client's account is not entering the target virtual space for the first time. When the client's account enters the target virtual space for the first time, the client can obtain the latency data corresponding to its operating system to obtain the client's initial device latency data. This initial device latency data can be stored internally by the client or on another device (e.g., stored on a server, from which the client obtains the initial device latency data). For example, for a client using Google's Android mobile operating system, its initial device latency data can be set to 250ms; for a client using Apple's iOS mobile operating system, its initial device latency data can be set to 80ms, etc.

[0114] In one feasible embodiment, for the first time a client's account enters the target virtual space, the above method may further include:

[0115] Once it is confirmed that the account is entering the target virtual space for the first time, obtain the client's operating system information and the initial device latency data corresponding to the client's operating system information.

[0116] Based on the initial device latency data, the background music and the voice data generated when the account first enters the target virtual space are aligned to obtain the target audio data.

[0117] In this embodiment, each time a client's account enters the target virtual space, the client can obtain the client's account identification information and match it with the identification information of accounts that have previously entered the target virtual space. If the match is successful, the client determines that the client's account is not entering the target virtual space for the first time; if the match is unsuccessful, the client determines that the client's account is entering the target virtual space for the first time. When the client determines that the client's account is entering the target virtual space for the first time, it can obtain the client's own operating system information and the initial device latency data corresponding to the client's own operating system information. Based on the initial device latency data, it aligns the background music with the voice data generated when the client's account first enters the target virtual space to obtain the target audio data. This allows for the acquisition of initial device latency data based on the client's operating system information when the client first enters the target virtual space. This initial device latency data is then used to determine historical device latency data, which is then fused with candidate device latency data obtained from subsequent visits to the target virtual space to obtain target device latency data for subsequent visits. This improves the accuracy of device latency determination, enabling more precise alignment of background music and voice data, enhancing the real-time communication effect of the target virtual space (e.g., the singing effect in an RTC karaoke room), and thus improving the client's account experience. Furthermore, the initial device latency data in this embodiment corresponds to the client's operating system, making it applicable to various manufacturers and types of client devices. Clients do not need to purchase various models of devices, reducing audio data processing costs and improving audio data processing efficiency.

[0118] Optionally, in step S11 above, if the client determines that its account has entered the target virtual space (e.g., not for the first time), the client can determine the playback device currently used to play the background music, i.e., what type of playback device the client is currently using to play the background music. This playback device can be an external speaker integrated into the client (e.g., a speakerphone) or an external device connected to the client (e.g., headphones connected to the client).

[0119] Optionally, in step S11 above, if the client determines that its account has entered the target virtual space (e.g., not for the first time), the client can also obtain the client's corresponding historical device latency data. This historical device latency data is the device latency data corresponding to the client's account's last entry into the target virtual space. It can be the client's initial device latency data or an update of the initial device latency data. For example, if the client's account is entering the target virtual space for the second time, the historical device latency data is the initial device latency data; if the client's account is entering the target virtual space for the third time, the historical device latency data is obtained by updating the initial device latency data; if the client's account is entering the target virtual space for the fourth time, the historical device latency data is obtained by updating the historical device latency data obtained from the third entry into the target virtual space.

[0120] Optionally, in step S11 above, if the client determines that the client's account has entered the target virtual space (e.g., not the first time entering the target virtual space), the client can also obtain the device type of the playback device, thereby obtaining the device delay determination information corresponding to the device type.

[0121] As an example, when the playback device is a speaker on the client, the client obtains echo cancellation delay determination information corresponding to the speaker; this echo cancellation delay determination information is determined to be device delay determination information. Specifically, the echo cancellation delay determination information is an alignment delay estimation technique based on Acoustic Echo Cancellation (AEC). In a target virtual space (e.g., an RTC karaoke room) where the client is performing through a speaker, the background music data being played is captured by the microphone, affecting the singing quality, and needs to be eliminated using AEC technology. In this embodiment, when the playback device is a speaker on the client, AEC technology is used to calculate the device delay determination information for the client. This achieves the use of corresponding device delay determination information based on the different device types of the playback device, further improving the accuracy of the client's device delay determination, enabling more precise alignment of background music and voice data, and further improving the generation accuracy of the target device delay data. Furthermore, since the echo cancellation delay determination information corresponds to the client's speaker, it can be applied to various manufacturers and types of client devices. Client accounts do not need to purchase various models of devices, reducing the cost of audio data processing and improving the efficiency of audio data processing.

[0122] As another example, when the playback device is a client-side headset, the client can obtain the baseband delay determination information corresponding to the headset; this baseband delay determination information is determined to be the device delay determination information. The baseband delay determination information is based on a delay calculation technique determined by baseband detection. When the client connects the headset to the target virtual space (e.g., an RTC karaoke room), the sound played from the headset is difficult to be captured by the microphone, and the AEC algorithm struggles to estimate the delay. Therefore, a different delay detection scheme is needed, namely, using an alignment delay estimation technique based on human voice baseband detection. In this embodiment, when the playback device is the client's headphones, the baseband delay meter information is used as the client's device delay determination information. This allows for the use of corresponding device delay determination information based on the type of playback device, further improving the accuracy of the client's device delay determination. This enables more precise alignment of background music and voice data, further enhancing the generation accuracy of the target device delay data. Furthermore, since the baseband delay determination information corresponds to the client's playback device, it can be applied to various manufacturers and types of client devices. The client account does not need to purchase various models of devices, reducing the cost of audio data processing and improving the efficiency of audio data processing.

[0123] In step S13, the delay between the background music and the voice data generated when the account enters the target virtual space is determined according to the device type of the playback device, and the candidate device delay data of the client in the target virtual space is obtained.

[0124] In this embodiment, since the client has pre-obtained the device type of the playback device, it can determine the delay between the background music and the voice data generated by the client's account when entering the target virtual space (e.g., not the first time entering the target virtual space) based on the device type of the playback device, thus obtaining the candidate device delay data of the client in the target virtual space. Optionally, the type of the voice data corresponds to the type of the target virtual space. In the case that the type of the target virtual space is an RTC karaoke room, the voice data can be songs sung by the client's account in the RTC karaoke room, etc.

[0125] In an optional embodiment, when the playback device is a speaker and the device delay determination information is echo cancellation delay determination information, step S13 may include:

[0126] Acquire reference data for background music and external audio data; external audio data refers to the sound data collected from the external environment, including the echo data generated by the background music played inside the client after being diffused through the speaker and the voice data of the client's account in the target virtual space; reference data is the raw data of the background music played inside the client.

[0127] Input reference data and external audio data into a linear filter. The linear filter filters out echo data in the external audio data to obtain filtered external audio data. Determine the delay between the reference data and the filtered external audio data to obtain candidate device delay data.

[0128] Figure 3 This is a schematic diagram illustrating the principle of echo cancellation delay determination information according to an exemplary embodiment, such as... Figure 3 As shown, the client-side (near-end) voice processing engine collects the raw data of the background music played by the client, i.e., the background music played by the client (i.e., the in-mix background music, which also serves as reference data), and also collects external audio data through the client's microphone. This external audio signal is the sound data collected from the external environment (including the in-mix background music reflected from the client's speaker to the microphone and the voice data emitted by the client in the target virtual space). The in-mix background music is played through the terminal's speaker, then propagates and reflects in the indoor environment before being collected by the terminal's audio acquisition module along with the voice data. Because the in-mix background music is an echo signal of the propagated and reflected in-mix background music, there is a certain delay between them. Without echo cancellation, the in-mix background music will be directly transmitted to the remote terminal along with the in-mix background music, causing a significant echo when the remote terminal receives the target audio data, resulting in poor sound quality.

[0129] As an example, the client can perform linear adaptive echo filtering on external audio data to eliminate the linear echo caused by externally mixed background music, resulting in filtered external audio data, thus removing the linear echo from the external audio signal. This echo filtering process is essentially a Finite Impulse Response (FIR) filter, used to simulate and generate an echo signal. Echo suppression is achieved by subtracting the simulated echo signal from the near-end signal. Its parameters can be adjusted based on the error between the predicted and actual echo signals. Once the filter parameters converge, the location of the first peak in the filter parameters is the device delay length.

[0130] As another example, in the AEC algorithm, device delay can also be detected through signal cross-correlation:

[0131] First, the signals in the near-end buffer and the far-end buffer are framed and then subjected to STFT (Short Time Fourier Transform) to obtain the near-end signal (Near) and the far-end signal (Far). The length of Far is determined by the empirical value of the client's maximum delay, denoted as n. To expand the delay search range, 2n data points will be placed in Far.

[0132]

[0133] Among them, Far * Let i represent conjugate. The i that makes R(i) reach its maximum value is the device delay.

[0134] As a third example, in the AEC algorithm, device delay can also be detected by signal coherence, calculated as follows:

[0135] Cov(Near(t),Far(ti))=a·Cov(Near(t-1),Far(ti))+(1-a)·Near(t)·Far * (ti)

[0136] Var(Near(t))=a·Var(Near(t-1))+(1-a)·Near(t)·Near * (t)

[0137] Var(Far(t))=a·Var(Far(t-1))+(1-a)·Far(t)·Far * (t)

[0138]

[0139] Where Cov() refers to the calculation of covariance, Far * and Near * Let represent conjugate, and 'a' be a hyperparameter, for example, it can be set to 0.99. The 'i' that makes C(i) take its maximum value is the device delay.

[0140] In this embodiment of the disclosure, when the playback device is a speaker, the client can use echo cancellation delay determination information to eliminate echo data in the external audio data. This prevents the external background music from being directly transmitted to the remote terminal along with the internal background music, which would cause a large echo when the remote terminal receives the target audio data, resulting in poor sound quality. This improves the accuracy of the target audio data determination. In addition, by comparing the frequency domain correlation between the reference data and the filtered external audio data, the delay data can be determined, which can improve the accuracy of the delay data determination, thereby improving the accuracy of further determining the target audio data.

[0141] In another alternative embodiment, Figure 4 This is a flowchart illustrating a method for determining candidate device latency data for a client in a target virtual space, according to an exemplary embodiment. Figure 4 As shown, when the playback device type is headphones and the device delay determination information is baseband delay determination information, the above step S13 may include:

[0142] In step S21, the music digital interface data of the background music is obtained; the music digital interface data includes the first pitch information sequence of the background music.

[0143] In step S23, the baseband frequency of the voice data of the account in the target virtual space is detected to obtain the baseband frequency detection result.

[0144] In step S25, the fundamental frequency detection result is subjected to pitch conversion processing to obtain the second pitch information sequence.

[0145] In step S27, the first pitch information in the first pitch information sequence is normalized to a preset scale to obtain the first target pitch information sequence, and the second pitch information sequence in the second pitch information sequence is normalized to a preset scale to obtain the second target pitch information sequence.

[0146] In step S29, offset information is determined when the similarity between the first target pitch information sequence and the second target pitch information sequence meets a preset condition, and candidate device delay data is obtained.

[0147] Optionally, in step S21 above, the Musical Instrument Digital Interface (MIDI) data is a descriptive musical language that uses digital control signals of notes to record music, including information such as each instrument, pitch, channel, duration, volume, and velocity. That is, the MIDI data for the background music may include the first pitch information sequence of the background music.

[0148] Optionally, in step S23 above, the client can use a preset baseband detection algorithm to perform baseband detection on the voice data of the account in the target virtual space, and obtain the baseband detection result, that is, obtain a human voice baseband f of a preset duration (e.g., 10ms).

[0149] Optionally, in step S25 above, the client can perform pitch conversion processing on the fundamental frequency detection result according to the following formula to obtain the second pitch information sequence (P), thereby converting the continuously input human voice data into the second pitch information sequence after segmenting it according to a preset duration:

[0150] P = 69 + 12 × log2(f / 440);

[0151] Where P is the second pitch information sequence and f is the fundamental frequency detection result. Optionally, considering the issue of musical octaves, the pitches need to be classified into 12 scales (c, c#, d, d#, e, f, f#, g, g#, a, a#, b). In step S27 above, the client can normalize the first pitch information in the first pitch information sequence to a preset scale according to the following formula to obtain the first target pitch information sequence, and normalize the second pitch information sequence in the second pitch information sequence to a preset scale to obtain the second target pitch information sequence:

[0152] P new = mod(pitch_information, 12) + 1;

[0153] Here, mod() refers to classifying pitch information into 12 musical notes, P new This refers to the normalized pitch information (either the first target pitch information sequence or the second target pitch information sequence).

[0154] Optionally, in step S29 above, the client can use the following formula to calculate the similarity between the first target pitch information sequence and the second target pitch information sequence, and the offset information when the preset conditions are met, to obtain the candidate device delay data:

[0155]

[0156] Among them, P newvocal This refers to the second target pitch information sequence, P newbgm This refers to the first target pitch information sequence, where i is a pitch in the pitch sequence, t is the offset between the two sequences, f(t) is the offset information, i.e., the candidate device delay data, and T is the upper limit of the empirical value for device delay (e.g., 1 second). n is the length of the audio data for a period of time (e.g., 5 seconds). When it < 0, P newbgm [it] is 0. The t that makes f(t) reach its minimum value, that is, the t when the similarity between the first target pitch information sequence and the second target pitch information sequence is the highest, is the candidate device delay data.

[0157] In this embodiment, when the playback device is a client-side headset, the fundamental frequency delay determination information is used as the client's device delay determination information. This effectively avoids the defect that the sound played by the headset is difficult to be captured by the microphone. It realizes the use of corresponding device delay determination information according to different playback device types, further improving the accuracy of client device delay determination. This allows background music and voice data to be more accurately aligned, thereby improving the generation accuracy of target audio data. In addition, during the fundamental frequency detection process, information such as musical octaves, pitch, and scale are fully considered to obtain a first target pitch information sequence and a second target pitch information sequence. Based on the similarity between the first target pitch information sequence and the second target pitch information sequence, the final candidate device delay data is determined, improving the determination accuracy of candidate device delay data and thus improving the generation accuracy of target audio data.

[0158] In a third alternative embodiment, Figure 5 This is a flowchart illustrating a method for determining candidate device latency data for a client in a target virtual space, according to an exemplary embodiment. Figure 5 As shown, step S13 above may further include:

[0159] In step S31, when the account's time in the target virtual space meets the preset time conditions, the time is processed in segments to obtain at least two time periods.

[0160] In this embodiment, if a client's account remains in the target virtual space for an extended period after entering it, the client can segment the time spent in a single instance of entering the target virtual space into at least two time periods. For example, the time spent in a single instance of entering the target virtual space can be divided into at least two time periods of 100 seconds each.

[0161] In step S33, candidate device latency data is obtained based on the latency between background music and account voice data in each time period, as well as the historical average device latency data corresponding to each time period.

[0162] Among them, the historical average device latency data corresponding to each time period represents the average latency information of the client's device latency data within the historical time period. The historical time period is the time period preceding each of at least two time periods; the latency between background music and account voice data within each time period is determined based on the device type of the playback device.

[0163] Optionally, in step S33 above, the client can determine the delay between background music and account voice data in each time period based on the device type of the device playing the audio, thus obtaining the client's device delay data in each time period. For specific calculations, please refer to the aforementioned baseband delay determination information and echo cancellation delay determination information, which will not be repeated here. After obtaining the client's device delay data in each time period, the client can fuse the device delay data in each time period with the historical average device delay data corresponding to each time period to obtain the final candidate device delay data. The historical average device delay data corresponding to each time period represents the average delay information of the client's device delay data within the historical time period. In this embodiment of the disclosure, when a client's account enters the target virtual space and remains there for an extended period, the time can be segmented. Based on the client's device latency data within each time segment and the historical average device latency data corresponding to each time segment, candidate device latency data is obtained. This results in the final candidate device latency data being the fusion of the device latency data within each time segment and the historical average device latency data corresponding to each time segment, thereby improving the generation accuracy of the candidate device latency data and consequently improving the generation accuracy of the target audio data.

[0164] In an exemplary embodiment, step S33 above may include:

[0165] Sort at least two time periods in ascending order to obtain a time period sequence.

[0166] Based on the device type of the playback device, the latency between the background music and the account's voice data in the first time period of the time sequence is determined, and the current historical average device latency data is obtained.

[0167] The current time period is determined by identifying the first remaining time period in the remaining time period sequence and then removing it from the remaining time period sequence. The remaining time period sequence is the sequence of time periods excluding the first remaining time period.

[0168] Based on the device type of the playback device, determine the delay between the background music and the account's voice data within the current time period, and obtain the current device delay data of the client within the current time period.

[0169] The current historical average device latency data and the current device latency data are merged, and the fusion result is redefined as the current historical average device latency data.

[0170] Repeat the process of determining the first remaining time period in the sequence of remaining time periods as the current time period, until the fusion result is redefined as the current historical average device latency data, until the current device latency data of the client in the last remaining time period of the sequence is fused to obtain candidate device latency data.

[0171] Step S33 above can be achieved using the following formula:

[0172] s1=(1-a)×s2+a×s last ;

[0173] Among them, s last This refers to the current device latency data, s1 refers to the current historical average device latency data, s2 refers to the previously calculated historical average device latency data, and 'a' is the weight information, which can be set according to actual business needs. For example, 'a' can be set to 0.1.

[0174] The following explanation uses at least two time periods, time period 1, time period 2, and time period 3, as an example to illustrate step S33.

[0175] The client sorts Time Period 1, Time Period 2, and Time Period 3 in ascending order according to time sequence, obtaining a time period sequence (e.g., Time Period 1 - Time Period 2 - Time Period 3). Based on the device type of the playback device, the client determines the latency between the background music and the account's audio data in the first time period (i.e., Time Period 1), obtaining the client's current historical average device latency data (i.e., s2) in Time Period 1.

[0176] The client takes the first time period (i.e., time period 2) in time period 2 and time period 3 as the current time period, and determines the delay between the background music and the account's voice data within the current time period (i.e., time period 2) based on the device type of the playback device, thus obtaining the client's current device delay data (i.e., s) within the current time period (i.e., time period 2). last ).

[0177] The client merges the current device latency data (i.e., s) within the current time period (i.e., time period 2) according to the above formula. last The client obtains the client's historical average device latency data (s2) within the current time period (i.e., within time period 2), and the client then uses the client's historical average device latency data within the current time period (i.e., within time period 2) as the current historical average device latency data.

[0178] The client takes time period 3 as the current time period and determines the delay between the background music and the account's voice data within the current time period (i.e., within time period 3) based on the device type of the playback device. This yields the client's current device delay data (i.e., s) within the current time period (i.e., within time period 3). last ).

[0179] The client merges the current device latency data (i.e., s) within the current time period (i.e., time period 3) according to the above formula. last ), and the client's current historical average device latency data in time period 2 (i.e., s2), to obtain the client's historical average device latency data in the current time period (i.e., time period 3) (i.e., s1).

[0180] After time period 3, the client's account logs out of the target virtual space, and the client uses the client's historical average device latency data within the current time period (i.e., within time period 3) as the final candidate device latency data.

[0181] In this embodiment, when a client's account enters the target virtual space and remains there for an extended period, the client can segment this time period and, in chronological order, first fuse the device latency data calculated in the second time period with the historical average device latency data of the first time period to obtain the historical average device latency data for the second time period. Then, the client can fuse the device latency data calculated in the third time period with the historical average device latency data for the second time period, and so on, until the client's account exits the target virtual space, thus obtaining the final candidate device latency data. This ensures that the final candidate device latency data is the result of fusing the device latency data in each time period with the corresponding historical average device latency data, improving the generation accuracy of the candidate device latency data and consequently improving the generation accuracy of the target audio data.

[0182] In one feasible embodiment, the device delay determination information includes echo cancellation delay determination information and baseband delay determination information. The current time period includes a first time period and a second time period. The playback device for the background music played by the client in the first time period is a speaker, and the playback device for the background music played by the client in the second time period is headphones.

[0183] Figure 6 This is a flowchart illustrating a method for determining current device latency data of a client within the current time period, according to an exemplary embodiment. Figure 6 As shown above, the delay between background music and account voice data within the current time period is determined based on the device type of the playback device, resulting in the client's current device delay data within the current time period. This can include:

[0184] In step S41, based on the echo cancellation delay determination information corresponding to the speaker, the delay between the background music and the account's voice data in the first time period is determined, and the first device delay data of the client in the first time period is obtained; based on the base frequency delay determination information corresponding to the headphones, the delay between the background music and the account's voice data in the second time period is determined, and the second device delay data of the client in the second time period is obtained.

[0185] In step S43, based on the time ratio information of the first time period and the second time period with the current time period, the delay data of the first device and the delay data of the second device are fused to obtain the delay data of the current device.

[0186] In this embodiment, if a client's account, during a certain time period (defined as the current time period) within the target virtual space, plays background music both through speakers and in headphone mode, then the device latency value (s) for that current time period is... last The calculation is as follows:

[0187]

[0188] Wherein, s1 is the delay data of the first device, s2 is the delay data of the second device, t1 is the duration of background music playback in speaker mode, and t2 is the duration of background music playback in headphone mode.

[0189] Optionally, in step S41 above, the client can determine the delay between background music and account voice data in the first time period based on the echo cancellation delay determination information, and obtain the client's first device delay data (s1) in the first time period; based on the baseband delay determination information, it can determine the delay between background music and account voice data in the second time period, and obtain the client's second device delay data (s2) in the second time period. In step S43 above, the client can calculate the first time ratio information between the first time period and the current time period, and calculate the second time ratio information between the second time period and the current time period. Then, it calculates the first product of the first time ratio information and the first device delay data (s1), and the second product of the second time ratio information and the second device delay data (s2); finally, it calculates the sum of the first product and the second product to obtain the current device delay data. This allows for the calculation of device latency when a client account enters the target virtual space for a specific time period (defined as the current time period), playing background music both through speakers and in headphone mode. The device latency is calculated using device latency determination information corresponding to different device types for different playback devices during the time period. By combining the device latency calculated using different device latency determination information, the accuracy of device latency data calculation is further improved, thereby improving the accuracy of target audio data generation.

[0190] In step S15, candidate device latency data and historical device latency data are merged to obtain the target device latency data of the client in the target virtual space.

[0191] In this embodiment, after obtaining the candidate device latency data, the client can fuse the candidate device latency data with the client's historical device latency data in the target virtual space to obtain the client's target device latency data in the target virtual space (for example, the client can fuse the candidate device latency data with the client's historical device latency data in the target virtual space before the first time to obtain the client's target device latency data in the target virtual space before the first time).

[0192] First, the process of generating historical device delay data will be introduced:

[0193] In one feasible implementation Figure 7 This is a flowchart illustrating a method for generating historical device delay data according to an exemplary embodiment, such as... Figure 7 As shown, the method for generating historical device delay data may include:

[0194] In step S51, when the account enters the target virtual space for the second time, the initial device latency data is determined to be the historical device latency data, and the candidate device latency data and the historical device latency data of the client in the target virtual space are merged to obtain the target device latency data.

[0195] In step S53, the historical device delay data is updated based on the target device delay data, and the updated device delay data is re-determined as the historical device delay data.

[0196] In step S55, when the account enters the target virtual space again, the process of repeatedly merging the candidate device latency data and historical device latency data currently in the target virtual space by the client continues until the updated device latency data is re-identified as historical device latency data, until the account exits the target virtual space.

[0197] In this embodiment, when the client's account enters the target virtual space for the second time, in step S51, the client can use the initial device latency data as historical device latency data, and fuse the client's current candidate device latency data in the target virtual space (i.e., the candidate device latency data for the second time in the target virtual space) with the historical device latency data (i.e., the initial device latency data) to obtain the target device latency data. In step S53, the client can use the target device latency data to update the historical device latency data, and re-determine the updated device latency data as the historical device latency data. In step S55, when the account enters the target virtual space again, the client repeats the above steps S51-S53 until the account exits the target virtual space.

[0198] The following example illustrates steps S51-S55:

[0199] Assume the client's account accesses the target virtual space four times:

[0200] When the client's account enters the target virtual space for the second time, in step S51 above, the client can use the initial device latency data as historical device latency data; and merge the client's current candidate device latency data in the target virtual space (i.e., the candidate device latency data for the second time in the target virtual space) with the historical device latency data (i.e., the initial device latency data) to obtain the client's target device latency data for the second time in the target virtual space. In step S53 above, the client updates the historical device latency data (i.e., the initial device latency data) based on the client's target device latency data for the second time in the target virtual space, and re-determines the updated device latency data as historical device latency data (i.e., the historical device latency data obtained after the second entry into the target virtual space).

[0201] In step S55 above, when the client's account enters the target virtual space for the third time, the client merges the candidate device latency data (i.e., the candidate device latency data for the third time in the target virtual space) and the historical device latency data (i.e., the historical device latency data obtained after the second entry into the target virtual space) to obtain the target device latency data for the third time in the target virtual space. Based on the target device latency data for the third time in the target virtual space, the client updates the historical device latency data (i.e., the historical device latency data obtained after the second entry into the target virtual space) and re-determines the updated device latency data as the historical device latency data (i.e., the historical device latency data obtained after the third entry into the target virtual space).

[0202] Similarly, when the client's account enters the target virtual space for the fourth time, the historical device delay data obtained after the client's account enters the target virtual space for the fourth time is determined in the same way as above, until the client's account exits the target virtual space and the process ends.

[0203] In this embodiment, each time a client's account enters the target virtual space, the existing historical device latency data is updated based on the target device latency data obtained by fusing the candidate device latency data and historical device latency data currently in the target virtual space. This updates the historical device latency data for the client's account entering the target virtual space, continuously improving the accuracy of historical device latency data determination. Furthermore, the continuous updating of historical device latency data ensures that the target device latency data obtained each time the client's account enters the target virtual space is obtained by fusing the latest historical device latency data, further improving the accuracy of target device latency data determination. This, in turn, improves the alignment accuracy of background music and voice data, and ultimately improves the generation accuracy of target audio data.

[0204] Optionally, in step S15 above, the client can use various methods to fuse candidate device latency data and historical device latency data to obtain the target device latency data of the client in the target virtual space, without making specific limitations on this.

[0205] In one implementation, the client can calculate the average of the candidate device latency data and the historical device latency data to obtain the target device latency data.

[0206] In another implementation, Figure 8 This is a flowchart illustrating a method for calculating delay data of a target device according to an exemplary embodiment, such as... Figure 8As shown, step S15 above may include:

[0207] In step S151, the weight information corresponding to the candidate device delay data and the historical device delay data is determined.

[0208] In step S153, the candidate device delay data and historical device delay data are weighted and aggregated according to the weight information to obtain the target device delay data.

[0209] In this embodiment, the process of steps S151-S153 described above can be achieved by the following formula:

[0210] t new = (1-μ)×t old +μ×s;

[0211] Where μ represents the first weight information of the candidate device latency data, s represents the candidate device latency data of the client in the target virtual space before the first time, (1-μ) represents the second weight information of the historical device latency data, and t old Delayed data for historical equipment.

[0212] In steps S151-S153 above, the first weight information of the candidate device latency data can be set according to actual business needs (for example, μ is set to 0.8). That is, the first weight information of the device latency data calculated when the client's account enters the target virtual space this time, and the difference between 1 and the first weight information is calculated to obtain the weight information of the historical device latency data. The client can calculate the product between the first weight information and the candidate device latency data, and the product between the second weight information and the historical device latency data, and calculate the sum of the two products to obtain the target device latency data of the client in the target virtual space before or after the first time. As can be seen, in this embodiment, when a client account accesses the target virtual space once, the device latency determination information is used to accurately calculate a candidate device latency data. By fusing this data with historical device latency data, the client's device latency can be made closer to the client's actual device latency, improving the accuracy of the client's device latency determination. This allows background music and voice data to be more accurately aligned, enhancing the real-time communication effect of the target virtual space (e.g., the singing effect in an RTC karaoke room), thereby improving the client's account experience. Furthermore, the target device latency data corresponds to the client's operating system and playback device type. Therefore, this embodiment can be applied to various manufacturers and types of client devices. The client account does not need to purchase various models of devices, reducing the cost of audio data processing and improving the efficiency of audio data processing.

[0213] In step S17, the background music and voice data are aligned according to the target device delay data to obtain the target audio data.

[0214] In this embodiment, after obtaining the target device delay data, the client can use various methods to align and overlay the background music and the voice data generated when the account enters the target virtual space for the first time to obtain the target audio data. This embodiment does not specifically limit this.

[0215] In one implementation, taking the target virtual space as an RTC karaoke room as an example, since the RTC karaoke room is real-time, the client can insert blank data corresponding to the target device's delay data at the beginning of the background music before using the microphone to collect voice data or before playing background music, so as to achieve alignment between the background music and voice data.

[0216] Figure 9 This is a flowchart illustrating another audio data generation method according to an exemplary embodiment, such as... Figure 9 As shown, taking the target virtual space as an RTC karaoke room as an example, the audio data generation method may include:

[0217] In step S61, the client's account enters the RTC karaoke room.

[0218] In step S63, the client determines whether the client's account is entering the RTC karaoke room for the first time.

[0219] In step S65, if the client's account is entering the RTC karaoke room for the first time, the client obtains the initial device latency data corresponding to the client's operating system. If the operating system is iOS, the initial device latency data corresponding to iOS is obtained; if the operating system is Android, the initial device latency data corresponding to Android is obtained.

[0220] In step S67, the background music and voice data are aligned based on the initial device delay data to obtain the target audio data.

[0221] In step S69, if the client's account is not entering the RTC karaoke room for the first time, the client determines whether the client is connected to headphones.

[0222] In step S611, if yes, the client obtains the baseband delay determination information corresponding to the headphones, and determines the delay between the background music and the voice data generated by the account entering the target virtual space for the first time based on the baseband delay determination information, thereby obtaining the candidate device delay data of the client in the target virtual space for the first time.

[0223] In step S613, if not, the client obtains the echo cancellation delay determination information corresponding to the speaker, determines the delay between the background music and the voice data generated when the account enters the target virtual space for the first time based on the echo cancellation delay determination information, and obtains the candidate device delay data of the client in the target virtual space for the first time.

[0224] In step S615, candidate device latency data and historical device latency data are merged to obtain target device latency data for the client in the target virtual space for the first time.

[0225] In step S617, the background music and voice data are aligned according to the target device delay data to obtain the target audio data.

[0226] Figure 10 This is a block diagram illustrating an audio data generation apparatus according to an exemplary embodiment. (Refer to...) Figure 10 The device includes a data acquisition module 71, a delay calculation module 73, a fusion module 75, and an alignment module 77.

[0227] The data acquisition module 71 is configured to acquire historical device latency data and the device type of the playback device when it is determined that the client's account has entered the target virtual space; the historical device latency data is the device latency data corresponding to the last time the account entered the target virtual space; the playback device is the playback device of the background music currently being played by the client.

[0228] The delay calculation module 73 is configured to determine the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtain the candidate device delay data of the client in the target virtual space;

[0229] The fusion module 75 is configured to merge candidate device delay data and historical device delay data to obtain the target device delay data of the client in the target virtual space;

[0230] Alignment module 77 is configured to perform alignment of background music and voice data based on target device delay data to obtain target audio data.

[0231] In an optional embodiment, the fusion module 75 includes:

[0232] The weight information determination unit is configured to determine the weight information corresponding to the candidate device delay data and the historical device delay data respectively;

[0233] The weight aggregation determination unit is configured to perform weight aggregation processing on candidate device delay data and historical device delay data based on weight information to obtain target device delay data.

[0234] In an optional embodiment, the apparatus further includes:

[0235] The initial device delay data acquisition module is configured to acquire the client's operating system and the initial device delay data corresponding to the client's operating system when it is determined that the account is entering the target virtual space for the first time.

[0236] The target audio data generation module is configured to perform actions based on the initial device delay data, aligning the background music and the voice data generated when the account first enters the target virtual space, to obtain the target audio data.

[0237] In an optional embodiment, the apparatus further includes a historical device delay data generation module, which includes:

[0238] The target device latency data generation unit is configured to, when an account enters the target virtual space for the second time, determine that the initial device latency data is historical device latency data, and merge the candidate device latency data of the client currently in the target virtual space with the historical device latency data to obtain the target device latency data;

[0239] The re-determination unit is configured to perform an update of historical device delay data based on target device delay data, and re-determine the updated device delay data as historical device delay data;

[0240] The repeated fusion unit is configured to perform the following operation when the account enters the target virtual space again: repeatedly fusion the candidate device latency data and historical device latency data of the client in the target virtual space, until the updated device latency data is re-identified as historical device latency data, until the account exits the target virtual space.

[0241] In an optional embodiment, the delay calculation module 73 includes:

[0242] The segmentation unit is configured to process time in segments to obtain at least two time periods when the account's time in the target virtual space meets a preset time condition.

[0243] The candidate device latency data generation unit is configured to perform a process based on the latency between the background music and the account's voice data in each time period, as well as the historical average device latency data corresponding to each time period, to obtain candidate device latency data.

[0244] Among them, the historical average device latency data corresponding to each time period represents the average latency information of the client's device latency data within the historical time period. The historical time period is the time period preceding each of at least two time periods; the latency between background music and account voice data within each time period is determined based on the device type of the playback device.

[0245] In an optional embodiment, the above-mentioned candidate device delay data generation unit includes:

[0246] The ascending sub-unit is configured to perform ascending sorting of at least two time periods to obtain a time period sequence;

[0247] The current historical average device latency data generation subunit is configured to determine the latency between the background music and the account's voice data in the time period sequence that is first sorted in the time period sequence based on the device type of the playback device, and obtain the current historical average device latency data.

[0248] The current time period determination sub-unit is configured to determine the first remaining time period in the remaining time period sequence as the current time period and delete the current time period from the remaining time period sequence; the remaining time period sequence is the sequence of time periods excluding the first remaining time period.

[0249] The current device delay data generation subunit is configured to determine the delay between the background music and the account's voice data in the current time period based on the device type of the playback device, and obtain the current device delay data of the client in the current time period;

[0250] The sub-unit is redefined and configured to perform the fusion of the current historical average device latency data and the current device latency data, and the fusion result is redefined as the current historical average device latency data;

[0251] The repeat execution subunit is configured to repeatedly determine the remaining time period at the top of the remaining time period sequence as the current time period, until the fusion result is re-determined as the current historical average device latency data, until the current device latency data of the client in the remaining time period at the bottom of the sequence is fused to obtain candidate device latency data.

[0252] In an optional embodiment, the current time period includes a first time period and a second time period. The playback device for the background music played by the client during the first time period is a speaker, and the playback device for the background music played by the client during the second time period is headphones.

[0253] The delay calculation module 73 includes:

[0254] The device delay data generation unit is configured to perform the following: determine the delay between background music and account voice data in a first time period based on the echo cancellation delay determination information corresponding to the speaker, and obtain the first device delay data of the client in the first time period; determine the delay between background music and account voice data in a second time period based on the base frequency delay determination information corresponding to the headphones, and obtain the second device delay data of the client in the second time period.

[0255] The delay data fusion unit is configured to perform the following operations: fused the delay data of the first device and the delay data of the second device based on the time ratio information of the first time period and the second time period with the current time period, respectively, to obtain the delay data of the current device.

[0256] In an optional embodiment, the delay calculation module 73 includes:

[0257] The music digital interface acquisition unit is configured to acquire music digital interface data of background music; the music digital interface data includes the first pitch information sequence of the background music.

[0258] The baseband detection unit is configured to perform baseband detection on the voice data of the account in the target virtual space and obtain the baseband detection result.

[0259] The pitch conversion unit is configured to perform pitch conversion processing on the fundamental frequency detection result to obtain a second pitch information sequence;

[0260] The normalization unit is configured to perform normalization processing on the first pitch information in the first pitch information sequence to a preset scale to obtain the first target pitch information sequence, and to perform normalization processing on the second pitch information sequence in the second pitch information sequence to a preset scale to obtain the second target pitch information sequence.

[0261] The offset information determination unit is configured to perform calculations on the similarity between the first target pitch information sequence and the second target pitch information sequence, and to obtain the offset information when the preset conditions are met, thereby obtaining candidate device delay data.

[0262] In an optional embodiment, the delay calculation module 73 includes:

[0263] The reference audio data acquisition unit is configured to acquire reference data and external audio data for background music; the external audio data is the sound data collected from the external environment, including the echo data generated by the background music played inside the client after being diffused through the speaker and the voice data of the account in the target virtual space; the reference data is the original data of the background music played inside the client.

[0264] The filtering unit is configured to input reference data and external audio data to a linear filter, filter echo data in the external audio data through the linear filter to obtain filtered external audio data, and determine the delay of the reference data and the filtered external audio data to obtain candidate device delay data.

[0265] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0266] In an exemplary embodiment, an electronic device for generating audio data is also provided, including a processor; a memory for storing processor-executable instructions; wherein, when the processor is configured to execute the instructions stored in the memory, it implements the steps of any of the audio data generation methods described above.

[0267] The electronic device can be a terminal, a server, or a similar computing device. Taking the electronic device as a client as an example... Figure 11 This is a block diagram illustrating an electronic device 80 for generating audio data according to an exemplary embodiment. The electronic device 80 can vary significantly due to different configurations or performance characteristics and may include one or more central processing units (CPUs) 81 (CPUs 81 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 83 for storing data, and one or more storage media 82 (e.g., one or more mass storage devices) for storing application programs 823 or data 822. The memory 83 and storage media 82 may be temporary or persistent storage. The program stored in the storage media 82 may include one or more modules, each module including a series of instruction operations on the electronic device. Furthermore, the CPU 81 may be configured to communicate with the storage media 82 and execute the series of instruction operations in the storage media 82 on the electronic device 80. Electronic device 80 may also include one or more power supplies 86, one or more wired or wireless network interfaces 85, one or more input / output interfaces 84, and / or one or more operating systems 821, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0268] The input / output interface 84 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 80. In one example, the input / output interface 84 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In an exemplary embodiment, the input / output interface 84 may be a radio frequency (RF) module for wireless communication with the Internet.

[0269] Those skilled in the art will understand that Figure 11 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, electronic device 80 may also include components that are more... Figure 11 The more or fewer components shown, or having the same Figure 11 The different configurations shown.

[0270] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the steps of any of the audio data generation methods described above.

[0271] In an exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the audio data generation method provided in any of the above embodiments.

[0272] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this disclosure can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0273] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0274] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for generating audio data, characterized in that, include: If it is determined that the client's account has not entered the target virtual space for the first time, historical device latency data and the device type of the playback device are obtained; the historical device latency data is the device latency data corresponding to the last time the account entered the target virtual space; the playback device is the playback device of the background music currently being played by the client; based on the device type of the playback device, the latency between the background music and the voice data generated by the account entering the target virtual space is determined to obtain the candidate device latency data of the client in the target virtual space; the candidate device latency data and the historical device latency data are fused to obtain the target device latency data of the client in the target virtual space; the background music and the voice data are aligned according to the target device latency data to obtain the target audio data; If it is determined that the account is entering the target virtual space for the first time, obtain the client's operating system information and the initial device latency data corresponding to the operating system information; Based on the initial device latency data, the background music and the voice data generated when the account first enters the target virtual space are aligned to obtain the target audio data; The method for generating historical device latency data includes: when the account enters the target virtual space for the second time, determining the initial device latency data as historical device latency data, fusing the candidate device latency data currently held by the client in the target virtual space with the historical device latency data to obtain target device latency data; updating the historical device latency data based on the target device latency data, and re-determining the updated device latency data as historical device latency data; when the account enters the target virtual space again, repeating the process of fusing the candidate device latency data currently held by the client in the target virtual space with the historical device latency data, up to the process of re-determining the updated device latency data as historical device latency data, until the account exits the target virtual space.

2. The audio data generation method according to claim 1, characterized in that, The process of fusing the candidate device latency data and the historical device latency data to obtain the target device latency data of the client in the target virtual space includes: Determine the weight information corresponding to the candidate device delay data and the historical device delay data respectively; Based on the weight information, the candidate device delay data and the historical device delay data are weighted and aggregated to obtain the target device delay data.

3. The audio data generation method according to claim 1, characterized in that, The step of determining the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtaining the candidate device delay data of the client in the target virtual space, includes: When the account's time in the target virtual space meets a preset time condition, the time is processed in segments to obtain at least two time periods; The candidate device latency data is obtained based on the latency between the background music and the voice data of the account in each time period, as well as the historical average device latency data corresponding to each time period. The historical average device latency data corresponding to each time period represents the average latency information of the client's device latency data within the historical time period, and the historical time period is the time period preceding each of the at least two time periods; the latency between the background music and the account's voice data within each time period is determined based on the device type of the playback device.

4. The audio data generation method according to claim 3, characterized in that, The step of obtaining the candidate device latency data based on the latency between the background music and the account's voice data in each time period, and the historical average device latency data corresponding to each time period, includes: Sort the at least two time periods in ascending order to obtain a time period sequence; Based on the device type of the playback device, determine the delay between the background music and the voice data of the account in the time period that is ranked first in the time period sequence, and obtain the current historical average device delay data; The remaining time period that is ranked first in the remaining time period sequence is determined as the current time period, and the current time period is deleted from the remaining time period sequence; the remaining time period sequence is the sequence of time periods excluding the time period ranked first. Based on the device type of the playback device, determine the delay between the background music and the voice data of the account in the current time period, and obtain the current device delay data of the client in the current time period; The current historical average device latency data and the current device latency data are merged, and the fusion result is re-determined as the current historical average device latency data; Repeat the process of determining the remaining time period that is first in the remaining time period sequence as the current time period, until the fusion result is re-determined as the current historical average device latency data, until the current device latency data of the client in the remaining time period at the end of the sequence is fused to obtain the candidate device latency data.

5. The audio data generation method according to claim 4, characterized in that, The current time period includes a first time period and a second time period. The client plays background music using a speaker during the first time period and headphones during the second time period. The step of determining the delay between the background music and the account's voice data within the current time period based on the device type of the playback device, and obtaining the client's current device delay data within the current time period, includes: Based on the echo cancellation delay determination information corresponding to the speaker, the delay between the background music and the voice data of the account in the first time period is determined, and the first device delay data of the client in the first time period is obtained; based on the base frequency delay determination information corresponding to the headphones, the delay between the background music and the voice data of the account in the second time period is determined, and the second device delay data of the client in the second time period is obtained. Based on the time ratio information of the first time period and the second time period to the current time period, the delay data of the first device and the delay data of the second device are fused to obtain the delay data of the current device.

6. The audio data generation method according to any one of claims 1 to 5, characterized in that, The step of determining the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtaining the candidate device delay data of the client in the target virtual space, includes: When the playback device is a headphone, the music digital interface data of the background music is acquired; the music digital interface data includes a first pitch information sequence of the background music. The fundamental frequency of the voice data of the account within the target virtual space is detected to obtain the fundamental frequency detection result. The fundamental frequency detection result is subjected to pitch conversion processing to obtain a second pitch information sequence; The first pitch information in the first pitch information sequence is normalized to a preset scale to obtain a first target pitch information sequence. The second pitch information sequence in the second pitch information sequence is normalized to the preset scale to obtain a second target pitch information sequence. The offset information is determined when the similarity between the first target pitch information sequence and the second target pitch information sequence meets a preset condition, and the candidate device delay data is obtained.

7. The audio data generation method according to any one of claims 1 to 5, characterized in that, The step of determining the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtaining the candidate device delay data of the client in the target virtual space, includes: When the playback device is a speaker, reference data for the background music and external audio data are acquired; the external audio data is the sound data collected from the external environment, including the echo data generated by the background music played inside the client after being diffused through the speaker and the voice data of the account in the target virtual space; the reference data is the original data of the background music played inside the client. The reference data and the external audio data are input into a linear filter, and the echo data in the external audio data is filtered by the linear filter to obtain the filtered external audio data; the delay of the reference data and the filtered external audio data is determined to obtain the candidate device delay data.

8. An audio data generation device, characterized in that, include: The data acquisition module is configured to acquire historical device latency data and the device type of the playback device when it is determined that the client's account has entered the target virtual space before; the historical device latency data is the device latency data corresponding to the last time the account entered the target virtual space; the playback device is the playback device of the background music currently being played by the client. The delay calculation module is configured to determine the delay between the background music and the voice data generated when the account enters the target virtual space based on the device type of the playback device, and obtain the candidate device delay data of the client in the target virtual space; The fusion module is configured to fuse the candidate device latency data and the historical device latency data to obtain the target device latency data of the client in the target virtual space; The alignment module is configured to perform alignment of the background music and the voice data based on the target device delay data to obtain target audio data; The initial device delay data acquisition module is configured to acquire the client's operating system information and the initial device delay data corresponding to the operating system information when it is determined that the account has entered the target virtual space for the first time. The target audio data generation module is configured to perform a process of aligning the background music and the voice data generated when the account first enters the target virtual space, based on the initial device delay data, to obtain target audio data. A historical device latency data generation module, comprising: a target device latency data generation unit, configured to, when the account enters the target virtual space for the second time, determine the initial device latency data as historical device latency data, and fuse the candidate device latency data currently held by the client in the target virtual space with the historical device latency data to obtain target device latency data; and a re-determination unit, configured to, update the historical device latency data based on the target device latency data, and re-determine the updated device latency data as historical device latency data. The repeat fusion unit is configured to perform the following operation when the account enters the target virtual space again: repeat the fusion of the client's current candidate device latency data and historical device latency data in the target virtual space until the updated device latency data is re-determined as historical device latency data, until the account exits the target virtual space.

9. An electronic device for generating audio data, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the audio data generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, wherein instructions in the computer-readable storage medium, when executed by a processor of an electronic device, cause the electronic device to perform the audio data generation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Audio track processing method and device, electronic equipment and storage medium

    CN113948054A