Audio playing method, system and device, electronic equipment and medium
Through the audio playback method of hardware link loopback, user audio data is collected and processed and directly transmitted to the speaker, solving the problem of excessive delay of wireless Bluetooth headphones, realizing a low-latency ear return experience, and improving the synchronization and audio quality of karaoke.
Patent Information
- Application Number
- CN202510449250.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-29
AI Technical Summary
The existing wireless Bluetooth headsets cannot meet the real-time requirements of less than 50 milliseconds in the Karoshima Ear Return scenario, resulting in the inability to align the user's voice with the accompaniment, affecting the singing experience.
User audio and background audio data are collected through the first microphone, and after noise reduction processing, the hardware link is used to loop back to the first speaker to play the ear return data to reduce software processing delay.
It achieves a low-latency ear return effect, better aligning the user's voice and accompaniment, improving the K-song experience and audio quality.
Smart Images

Figure CN120386507A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data processing, and particularly relates to an audio playback method, system, device, electronic device, and medium. Background Art
[0002] With the gradual development of live broadcast applications, short video applications, and singing applications, users often perform online dubbing or online karaoke through application functions and publish their works. Currently, the following technical bottlenecks are encountered:
[0003] On the one hand, with the rise of wireless earphones, wired earphones have gradually been phased out by the market due to poor portability. On the other hand, in-ear monitors require very low latency, for example, less than 50 milliseconds (ms), otherwise it is difficult for the human ear to align with the accompaniment, and the real-time performance of wireless Bluetooth transmission cannot meet this requirement. Although wireless earphones have obvious advantages in other aspects, they cannot provide a stable and low-latency real-time in-ear monitor experience.
[0004] Therefore, it is currently difficult to meet users' demands for real-time in-ear monitors. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide an audio playback method, system, device, electronic device, and medium that can meet users' demands for real-time in-ear monitors.
[0006] In a first aspect, the embodiments of this application provide an audio playback method, which includes:
[0007] Collect user audio data and background audio data through a first microphone to obtain first audio data;
[0008] Process the first audio data to obtain in-ear monitor data;
[0009] Play the in-ear monitor data through a first speaker;
[0010] Among them, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
[0011] In a second aspect, the embodiments of this application provide an audio playback system, which includes:
[0012] A first microphone, configured to collect user audio data and background audio data to obtain first audio data;
[0013] A real-time in-ear monitor module, configured to process the first audio data to obtain in-ear monitor data;
[0014] A first speaker, configured to play the in-ear monitor data;
[0015] Among them, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
[0016] In a third aspect, an embodiment of the present application provides an audio playback device, which includes:
[0017] An acquisition module, configured to acquire user audio data and background audio data through a first microphone to obtain first audio data;
[0018] A processing module, configured to process the first audio data to obtain in-ear monitor data;
[0019] A playback module, configured to play the in-ear monitor data through a first speaker;
[0020] Among them, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
[0021] In a fourth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0022] In a fifth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0023] In a sixth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.
[0024] In a seventh aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.
[0025] In the embodiments of the present application, user audio data and background audio data are acquired through a first microphone to obtain first audio data, providing a complete data basis for subsequent audio processing, so as to accurately feedback the user's voice to the user. The first audio data is processed to obtain in-ear monitor data, which can remove the noise components in the first audio data and improve the audio quality. The in-ear monitor data is played through a first speaker. Since the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link, the audio transmission delay is reduced through the hardware link, enabling the user to hear their own voice better aligned with the background audio data and meeting the user's demand for real-time in-ear monitoring. Description of the Drawings
[0026] Figure 1 is a flowchart of an audio playback method provided by an embodiment of the present application;
[0027] Figure 2 is a schematic diagram of a microphone and a speaker in an electronic device provided by an embodiment of the present application;
[0028] Figure 3 is a schematic diagram of a data stream of audio data provided by an embodiment of the present application;
[0029] Figure 4 is a structural diagram of an audio playback device provided by an embodiment of the present application;
[0030] Figure 5 is one of the schematic diagrams of the hardware structure of an electronic device provided by an embodiment of the present application;
[0031] Figure 6 is the second schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0032] Next, the technical solutions of the embodiments of the present application will be clearly described in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0033] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0034] The information display method provided by the embodiments of the present application can be applied to at least the following application scenarios, which will be described below.
[0035] Currently, when using karaoke apps, users often want to use the in-ear monitoring function. This allows singers to hear their own voice and the accompanying music through headphones, allowing them to better grasp pitch and rhythm. To ensure that the singer's voice and accompaniment are synchronized and achieve a good singing experience, the industry generally requires the in-ear monitoring delay to be less than 50 milliseconds. If the delay exceeds this standard, the singer will notice a noticeable delay between their voice and the accompaniment, resulting in a misalignment, which will affect the singer's performance and experience.
[0036] Wireless Bluetooth headphones use Bluetooth technology to transmit audio signals. However, Bluetooth transmission can sometimes experience a certain delay due to the inherent characteristics of Bluetooth technology and various factors during signal transmission. This delay can exceed the 50 milliseconds required for in-ear karaoke monitoring. Therefore, the real-time performance of wireless Bluetooth transmission is insufficient to meet the needs of in-ear karaoke monitoring scenarios. This means that when using wireless Bluetooth headphones for karaoke monitoring, the sound heard by the human ear may not align with the accompaniment.
[0037] In response to the problems arising from related technologies, the embodiments of the present application provide an audio playback method, system, device, electronic device and medium, which can solve the problem in related technologies that it is difficult to meet users' needs for real-time ear monitoring.
[0038] The audio playback method provided in the embodiment of the present application is described in detail below with reference to specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0039] Figure 1 A flowchart of an audio playback method provided in an embodiment of the present application.
[0040] like Figure 1 As shown, the audio playback method may include steps 110 to 130, and the method is applied to an audio playback device, as shown below:
[0041] Step 110: collecting user audio data and background audio data through a first microphone to obtain first audio data;
[0042] The first microphone is a device used to collect sound data. For example, in a karaoke scene, it is responsible for capturing the sound made by the user when singing, that is, the user audio data; and the sound in the surrounding environment, that is, the background audio data, providing raw materials for subsequent audio processing.
[0043] The first speaker is used to play the processed in-ear feedback data, allowing users to hear their own voice feedback in real time to achieve the in-ear feedback effect.
[0044] The first microphone uses the electroacoustic conversion element inside it to convert the sound vibrations in the air into electrical signals. These electrical signals represent the user audio data and background audio data, thus obtaining the digitized first audio data. By comprehensively collecting the user's voice information, it provides a complete data basis for subsequent ear monitoring and audio processing, so as to accurately feedback the user's singing voice to the user.
[0045] Step 120: Process the first audio data to obtain ear monitoring data;
[0046] The collected first audio data needs to go through a series of processes to obtain ear monitoring data. The processing process can include: performing noise reduction processing to remove the noise components in the background audio data through algorithms and improve the clarity of the sound; then audio enhancement can be performed on the user audio data, such as increasing the volume, adjusting the tone color, adding reverberation effects, etc., to improve the quality and expressiveness of the sound. After these processes, the obtained is the ear monitoring data.
[0047] The ear monitoring data enables the user to hear their own optimized voice in real time, so that the user can monitor and adjust their vocalization. For example, when giving a speech, the user can better control the volume and rhythm through ear monitoring, or when conducting voice training, they can promptly discover inaccurate pronunciations and correct them.
[0048] The following is illustrated in combination with two specific scenarios:
[0049] Scenario 1: When the user is using a KTV application, the user usually uses a microphone to sing songs. The first microphone will simultaneously collect the user audio data and the background sounds in the surrounding environment. These collected sound signals combined together form the first audio data.
[0050] Process the first audio data, that is, separate the user's singing voice from the first audio data and perform optimization processing on it, such as adjusting effects like volume, tone color, and reverberation, and then send the processed singing voice as ear monitoring data to the headphones worn by the user. In this way, when singing, the user can hear their own processed voice through the headphones, so as to better grasp the pitch, rhythm, and singing effect, and at the same time avoid being unable to clearly hear their own singing voice due to factors such as environmental noise.
[0051] Scenario 2: When the user is using a dubbing application, the user needs to perform voice performances according to the characters and plots in materials such as videos or animations. The first microphone will collect the user audio data and the background sounds in the recording environment, jointly constituting the first audio data. For example, when the user is dubbing a character in an animated film, the microphone will collect the user's voice imitating the character and the slight environmental noise in the recording studio, etc., to form the first audio data.
[0052] In the dubbing scenario, in-ear monitor data is equally important. Processing the first audio data is mainly to enable users to hear their own dubbed voices more clearly so that they can adjust their performances in a timely manner. The processing may include removing background noise, enhancing the clarity and expressiveness of the voice, etc., and then transmitting the processed user-dubbed voice as in-ear monitor data to the user's headphones. In this way, users can hear their own voices in real-time through the in-ear monitor when dubbing, and more accurately adjust the intonation, speech rate, tone, etc. of the voice according to the emotions of the character and the needs of the scene to achieve a better dubbing effect.
[0053] Step 130, play the in-ear monitor data through the first speaker;
[0054] Among them, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
[0055] Hardware link loopback: It means that after the audio data is collected by the first microphone, instead of going through multiple transfers and processes at the complex software level, it is directly transmitted to the first speaker through a specific signal line or connection method at the hardware level. This method reduces the delay caused by software processing and can improve the real-time performance of data transmission.
[0056] The first audio data is directly looped back to the first speaker through the hardware link, and finally the first speaker converts the electrical signal into a sound signal for playback, forming the in-ear monitor data.
[0057] Through the hardware link loopback, the audio transmission delay is greatly reduced. Since the complex software processing process is avoided, the playback delay of the in-ear monitor data can be made as low as possible, meeting the standard that the in-ear monitor for KTV requires a delay of less than 50 ms, enabling users to better align their own voices with the accompaniment and enhancing the KTV experience.
[0058] In a possible embodiment, the following steps may further be included:
[0059] Step 140, play the background audio data through the second speaker;
[0060] Step 150, collect the user audio data and the background audio data through the second microphone to obtain second audio data;
[0061] Step 160, perform noise reduction processing on the first audio data to obtain the in-ear monitor data;
[0062] Step 170, perform noise reduction processing on the second audio data to obtain the processed second audio data;
[0063] Step 180, perform fusion processing on the in-ear monitor data, the processed second audio data, and the background audio data to obtain target audio data.
[0064] Second speaker: mainly used for specifically playing background audio data, such as the accompaniment music during karaoke singing, to create a suitable music atmosphere.
[0065] Second microphone: used to collect sound signals and work together with the first microphone to collect user audio data and background audio data in multiple directions, obtaining more comprehensive sound information.
[0066] The following explains each step in sequence;
[0067] Step 140: After the second speaker receives the signal of the background audio data, through the internal electro-acoustic conversion element, it converts the electrical signal into a sound signal and plays it out, allowing the user to hear the background sound.
[0068] Playing the background audio data through the second speaker alone can present the background sound more clearly and independently, enhancing the user's music experience during karaoke singing and allowing the user to better immerse in the music atmosphere.
[0069] Step 150: The second microphone uses the principle of electro-acoustic conversion to convert the surrounding sound vibrations into electrical signals, thereby collecting user audio data and background audio data to form second audio data.
[0070] Adding the second microphone to collect sound can obtain richer and more comprehensive sound information, make up for the possible acquisition blind spots of the first microphone, and improve the integrity and accuracy of the audio data.
[0071] Step 160: Analyze and process the first audio data using a noise reduction algorithm. The algorithm will identify the noise characteristics in the audio and then remove the noise through means such as filtering to obtain relatively pure in-ear monitor data.
[0072] Noise reduction processing: refers to using specific algorithms or technologies to remove noise interference in audio data. In the karaoke scene, the noise may come from environmental noise, etc. Noise reduction processing can make the audio cleaner and clearer.
[0073] Performing noise reduction processing on the first audio data can remove the noise in the in-ear monitor data, making the user's own voice heard more pure and clear, and enhancing the in-ear monitor effect.
[0074] In a possible embodiment, step 160 may specifically include the following steps:
[0075] Perform howling suppression on the first audio data according to the first howling processing parameter, the first link delay, and the user voice feature to obtain the first audio data after howling suppression; the first link delay is used to describe the link delay of the user audio data from being played by the first speaker to being collected by the first microphone;
[0076] Perform echo suppression processing on the first audio data after howling processing according to the first echo processing parameter to obtain in-ear monitor data.
[0077] First link delay: A specific time interval, that is, the time elapsed from when the user audio data is played from the first speaker to when it is collected by the first microphone. This delay is caused by the transmission and processing of audio in hardware devices, transmission lines, etc., and has an important impact on the real-time in-ear monitor effect. Because if the delay is too large, the in-ear monitor sound heard by the user will be out of sync with the sound they actually emit.
[0078] User voice characteristics: Parameters extracted from user audio data that can represent the characteristics of the user's voice, such as pitch, timbre, speech rate, intonation, etc. These characteristics can be used to distinguish the user's voice from other interfering sounds and help process audio data more accurately.
[0079] First howling processing parameter: A series of parameters used to process howling problems in audio. Howling refers to the phenomenon in an audio system where, due to sound feedback forming a positive feedback loop, the audio signal is continuously amplified, resulting in a sharp and ear-piercing sound. These parameters can guide the algorithm on how to identify and eliminate howling.
[0080] First echo processing parameter: Parameters used for echo suppression processing. Echo refers to the repeated sound formed when sound reflects off an obstacle during propagation and is superimposed on the original sound. These parameters help the algorithm determine how to identify and eliminate echoes.
[0081] First audio data after howling processing: The audio data obtained after performing howling processing on the first audio data, at which time the howling problem in the audio has been solved.
[0082] First, record the process of the user audio data being played from the first speaker to being collected by the first microphone, so as to measure the first link delay. Through the audio feature extraction algorithm, extract the user voice characteristics from the user audio data, such as using methods like spectrum analysis and pitch detection. The first howling processing parameter and the first echo processing parameter can be pre-set during system initialization or in previous test processes, or may be dynamically adjusted according to the real-time audio situation.
[0083] Among them, in the step of suppressing howling of the first audio data according to the first howling processing parameter, the first link delay, and the user voice characteristics, the following steps may specifically be included:
[0084] Remove the first human voice interference data in the first audio data according to the first howling processing parameter, the first link delay, and the user voice characteristics to obtain the first audio data after howling processing.
[0085] The first human voice interference data refers to the interference of other human voices in the first audio data, excluding the user's normal singing voice, such as the noisy voices of people in the surrounding environment, the voices of other people mixed in the audio system, etc.
[0086] Combined with the first link delay, considering the time factors of sound propagation and feedback, the position where the howling occurs can be more accurately located. Using the user's voice characteristics, the sounds in the first audio data are compared with the user's voice characteristics to distinguish the user's normal singing voice and the first human voice interference data. The first human voice interference data is removed through methods such as filtering and noise reduction to obtain the first audio data after howling processing.
[0087] According to the first echo processing parameter, the possible echo components in the first audio data after howling processing are analyzed. Through technologies such as adaptive filtering, the echo is estimated and eliminated, so that the echo problem in the finally obtained in-ear monitor data is solved.
[0088] Accurately obtaining the first link delay helps to consider the time factor of sound propagation in subsequent processing, avoid the problem of out-of-sync sound caused by delay, and improve the real-time performance and accuracy of the in-ear monitor. Extracting the user's voice characteristics can better distinguish the user's voice from other interfering sounds, provide a basis for subsequent removal of human voice interference and audio processing, and make the processed audio more purely retain the user's singing voice. The first howling processing parameter and the first echo processing parameter provide specific operation bases for howling processing and echo suppression, ensuring that the howling and echo problems in the audio can be effectively solved.
[0089] Therefore, removing the first human voice interference data can make the in-ear monitor data cleaner, reduce the interference of environmental noise and other human voices, make the user hear their own voice more clearly, and improve the KTV experience. Solving the howling problem can avoid the interference of sharp and ear-piercing howling sounds to the user. After the echo suppression processing, there is no obvious echo in the in-ear monitor data, and the sound heard by the user is more real and natural, without the phenomenon of sound tailing and repetition, further improving the quality of the in-ear monitor.
[0090] Among them, in the step of performing echo suppression processing on the first audio data after howling processing according to the first echo processing parameter to obtain the in-ear monitor data, it may specifically include the following steps:
[0091] Determine the first signal relationship between the background audio data output by the second speaker and the input signal of the first microphone;
[0092] Predict the first echo audio data according to the first signal relationship;
[0093] Obtain the in-ear monitor data according to the first echo processing parameter and the first echo audio data in the first audio data after removing howling processing.
[0094] First signal relationship: It refers to the correlation characteristics between the background audio data output by the second speaker and the input signal of the first microphone, including the relationships in aspects such as the amplitude, phase, and time delay of the signal. This relationship reflects how the background audio is collected by the first microphone during propagation and its interaction with other sound signals.
[0095] First echo audio data: It refers to the echo signal generated when the background audio output by the second speaker is reflected in space and then collected again by the first microphone. This part of the signal will interfere with the user's normal audio collection and needs to be suppressed.
[0096] Analyze the background audio data output by the second speaker and the input signal collected by the first microphone. Methods such as correlation analysis and spectrum analysis can be used to find the corresponding relationships between the two in terms of amplitude, phase, and time. For example, by calculating the cross-correlation function of the two signals, determine the time delay between them to understand how long it takes for the background audio to be collected by the first microphone; by comparing the frequency components of the two signals through spectrum analysis, analyze the amplitude change situation.
[0097] After determining the first signal relationship, use the first signal relationship to predict the first echo audio data. Since the echo is formed after the background audio is reflected, its characteristics are somewhat related to the original background audio. According to the previously obtained signal relationships, such as time delay and amplitude attenuation information, perform corresponding transformations and processing on the background audio to simulate the echo audio signal collected at the first microphone.
[0098] According to the first echo processing parameters, these parameters specify the specific algorithms and strategies for echo suppression, such as using an adaptive filtering algorithm, etc. Subtract the predicted first echo audio data from the first audio data after feedback suppression to achieve echo suppression. In this way, the echo interference in the audio is removed, and pure in-ear monitor data is obtained, making the sound heard by the user clearer and more real.
[0099] Accurately determining the first signal relationship provides a basis for subsequent prediction of echo audio data. Only by clearly understanding the correlation between the background audio and the input signal of the first microphone can the characteristics of the echo signal be predicted more accurately, improving the effect of echo suppression. Being able to relatively accurately predict the first echo audio data makes the subsequent echo suppression processing have a clear goal. By predicting the echo signal in advance, it avoids blindly performing filtering or attenuation operations in actual processing, improving the pertinence and efficiency of the processing.
[0100] As a result, the echo interference in the audio is effectively removed, making the in-ear monitor data purer. When users use the in-ear monitor function, they will not be affected by echoes, and the sound of their own voices they hear is clearer and more natural, greatly enhancing the KTV experience and audio quality. At the same time, it also helps to further process and save the audio data subsequently, improving the quality of the final audio work.
[0101] Step 170: Process the second audio data using a noise reduction algorithm to identify and remove the noise therein, obtaining the processed second audio data.
[0102] Perform noise reduction processing on the second audio data to make the processed second audio data cleaner, providing high-quality audio material for subsequent fusion processing.
[0103] In a possible embodiment, step 170 may specifically include the following steps:
[0104] Perform howling suppression on the second audio data according to the second howling processing parameter, the second link delay, and the user voice feature, obtaining the second audio data after howling processing; the second link delay is used to describe the link delay of the user audio data from being played by the second speaker to being collected by the second microphone;
[0105] Perform echo suppression processing on the second audio data after howling processing according to the second echo processing parameter, obtaining the processed second audio data.
[0106] Second link delay: Used to measure the time interval experienced by the user audio data from being played by the second speaker to being collected by the second microphone, reflecting the time delay situation of the audio signal on the entire path from being emitted by the second speaker and propagated through space to being received by the second microphone, which is crucial for understanding the time characteristics of audio signal transmission and processing.
[0107] User voice feature: A series of parameters or feature values extracted from the user audio data that can characterize the user's voice characteristics, such as fundamental frequency, formant position, rhythm and prosody of speech, etc. These features can be used to distinguish the user's voice signal from other interference signals.
[0108] Second howling processing parameter: A parameter specifically set for the possible howling phenomenon in the second audio data, which can guide the algorithm to identify and process the howling signal to eliminate or reduce the impact of howling on audio quality.
[0109] Second echo processing parameter: Relevant parameters for processing echo problems in the second audio data, such as echo attenuation coefficient, delay time adjustment parameter, etc. Through these parameters, the algorithm can effectively suppress and eliminate the echo signal, making the audio clearer.
[0110] Among them, in the step of suppressing howling in the second audio data according to the second howling processing parameter, the second link delay, and the user voice feature to obtain the second audio data after howling processing, the following steps may specifically be included:
[0111] Remove the second human voice interference data in the second audio data according to the second howling processing parameter, the second link delay, and the user voice feature to obtain the second audio data after howling processing.
[0112] Second human voice interference data: The other human voice components existing in the second audio data except for the voice content that the user expects to retain, which may be the voices of other people in the environment, a noisy human voice background, etc. These interference data will affect the purity and quality of the audio.
[0113] Second audio data after howling processing: The audio data obtained after performing howling processing on the second audio data. At this time, the howling phenomenon in it has been effectively suppressed or eliminated, and the audio quality has been improved to a certain extent.
[0114] Play the user audio data and record the time difference from the emission from the second speaker to the collection by the second microphone, so as to obtain the second link delay. Use a specific audio feature extraction algorithm to analyze and extract the user voice feature from the user audio data. For example, adopt a spectrum analysis method based on Fourier transform to obtain the frequency feature of the voice, or extract the parameter feature of the voice through linear predictive coding.
[0115] The second howling processing parameter and the second echo processing parameter can be preset empirical values, or can be dynamically adjusted by the system according to the previous audio processing situation or a specific learning algorithm. According to the second howling processing parameter, the algorithm scans the second audio data to identify the possible howling frequencies and characteristic patterns. At the same time, in combination with the second link delay, consider the time factor of the audio signal propagating and feeding back in space to more accurately locate the position where the howling occurs.
[0116] Utilize the extracted user voice feature to compare and match the sound components in the second audio data with the user voice feature. Through methods such as similarity calculation, distinguish the user's voice signal and the second human voice interference data, and then use technical means such as filtering and noise reduction to remove these interference data to obtain the second audio data after howling processing.
[0117] According to the second echo processing parameter, the algorithm analyzes the second audio data after howling processing to detect possible echo signals therein. According to the characteristics of the echo, such as the delay time, amplitude size, etc., algorithms such as adaptive filtering and echo cancellation are used to estimate and eliminate the echo signal, so as to obtain the processed second audio data, and there is no obvious echo interference in the audio.
[0118] Accurately obtaining the second link delay helps to accurately consider the time factor in subsequent audio processing, avoid the problem of audio signal asynchronization caused by the delay, and ensure the real-time and accuracy of audio processing. Extracting user voice features provides a key basis for distinguishing user voice and interference signals, enabling more accurate retention of the user's voice content when removing interference data, and improving the purity and intelligibility of the audio.
[0119] Through the second howling processing parameter and the second echo processing parameter, it is possible to more effectively cope with the howling and echo problems that may occur in the audio, and improve the audio quality. Successfully removing the second human voice interference data makes the second audio data purer, reduces the impact of external human voice interference on the user's voice, makes the user's voice content more prominent and clear, and improves the quality and audibility of the audio.
[0120] The processing of howling effectively eliminates the sharp and harsh sounds in the audio and also improves the overall audio environment. After the echo suppression processing, there is no obvious echo phenomenon in the processed second audio data, and the audio is clearer and more natural, so that the user will not be interfered by the echo when listening to or using the audio data, improving the user experience and providing a better basis for subsequent audio processing and applications.
[0121] Among them, in the step of performing echo suppression processing on the second audio data after howling processing according to the second echo processing parameter to obtain the processed second audio data, the following steps may specifically be included:
[0122] Determine the second signal relationship between the background audio data output by the second speaker and the input signal of the second microphone;
[0123] Predict the second echo audio data according to the second signal relationship;
[0124] Remove the second echo audio data in the second audio data according to the second echo processing parameter to obtain the processed second audio data.
[0125] Second signal relationship: Describes the inherent connection between the background audio data output by the second speaker and the input signal of the second microphone, covering various correlation characteristics such as the amplitude, phase, and time delay of the signal. By analyzing this relationship, it is possible to understand how the background audio is received by the second microphone during propagation and its interaction with other sound components.
[0126] Second echo audio data: The part of the audio signal that is collected by the second microphone again after the background audio output by the second speaker is reflected by an obstacle during propagation. The echo will interfere with the original audio signal, resulting in a blurred and unclear sound, so echo suppression processing is required.
[0127] Perform a detailed analysis of the background audio data output by the second speaker and the input signal collected by the second microphone. Signal processing techniques such as correlation analysis and spectral analysis are usually adopted to calculate parameters such as the correlation and frequency response between the two signals. For example, the cross-correlation function is calculated to determine the time delay between the two signals, and spectral analysis is used to compare their frequency components and amplitude distributions, so as to comprehensively determine the second signal relationship.
[0128] After clarifying the second signal relationship, use the second signal relationship to predict the second echo audio data. Since the echo is formed by the reflection of the background audio, its characteristics are related to the original background audio to a certain extent. According to the previously determined signal relationship, such as time delay, amplitude attenuation, etc., perform corresponding transformations and processing on the background audio to simulate the echo audio signal that may be collected at the second microphone. For example, if it is determined that the echo has a certain time delay and amplitude attenuation, the background audio can be delayed and attenuated to predict the echo.
[0129] According to the second echo processing parameters, specific algorithms and strategies for echo suppression are specified, such as the parameters of the adaptive filtering algorithm, filter coefficients, etc. Subtract the predicted second echo audio data from the second audio data after howling suppression to achieve echo suppression. By continuously adjusting the filter coefficients, the algorithm can adaptively track the changes of the echo, so as to more effectively remove the echo interference and obtain the processed second audio data.
[0130] Accurately determining the second signal relationship is the basis for subsequent echo prediction and suppression. It provides key information for predicting the second echo audio data, making the prediction results more accurate and reliable. Only by clearly understanding the relationship between the background audio and the input signal of the second microphone can the characteristics of the echo be more accurately simulated, providing strong support for echo suppression.
[0131] It can accurately predict the second voice audio data, enabling clear targets for echo cancellation processing. By predicting the echo signal in advance, it avoids blindly performing filtering or attenuation operations during actual processing, improving the pertinence and efficiency of processing. This helps to maximize the removal of echo interference without sacrificing the quality of the original audio signal.
[0132] It effectively removes the echo interference in the second audio data, making the processed second audio data purer and clearer. When users use this audio data, they will not be affected by echoes and will hear more real and natural sounds. This is of great significance for enhancing the KTV experience, improving audio quality, and subsequent audio processing and applications.
[0133] Among them, the first speaker and the first microphone are arranged on the first side of the electronic device, and the second speaker and the second microphone are arranged on the second side of the electronic device;
[0134] The second howling processing parameter is less than the first howling processing parameter corresponding to the first audio data;
[0135] The second echo processing parameter is greater than the first echo processing parameter corresponding to the first audio data.
[0136] As Figure 2 shown, the first speaker 22 and the first microphone 21 are arranged on the first side of the electronic device; the second speaker 32 and the second microphone 31 are arranged on the second side of the electronic device;
[0137] It can be understood that it is also possible to arrange the first speaker 32 and the first microphone 31 on the first side of the electronic device; the second speaker 22 and the second microphone 21 are arranged on the second side of the electronic device.
[0138] The two microphones are on different sides of the electronic device respectively, and can collect sounds from different angles, which helps to obtain more stereo and richer audio data and reduce the blind area of sound collection. For example, in a KTV scenario, it can better capture the sound characteristics of the user's voice and the surrounding environment.
[0139] The two speakers play audio on different sides of the electronic device respectively, which can create a broader sound field effect. For example, when playing background audio, it can let users feel a more immersive music atmosphere and enhance the KTV experience.
[0140] Among them, the first howling processing parameter and the second howling processing parameter are used to describe the intensity of howling suppression processing. Howling is usually caused by positive feedback of sound. The first speaker and the first microphone are relatively close, and it is easier for sound to form a feedback loop, resulting in a greater possibility of howling. The second speaker and the second microphone are located on the second side of the electronic device, with a relatively longer distance, and the possibility of sound feedback forming howling is relatively small. Therefore, when processing the second audio data, it is not necessary to set a relatively high howling processing parameter as when processing the first audio data.
[0141] Due to the close distance between the first speaker and the first microphone, the signal intensity of the sound emitted by the speaker received by the microphone is relatively large, and it is easier to generate howling interference. The signal intensity of the sound of the second speaker received by the second microphone is relatively weak, and the degree of howling interference is also relatively low. Therefore, the second howling processing parameter can be set to be relatively small.
[0142] Among them, the first echo processing parameter and the second echo processing parameter are used to describe the intensity of echo suppression processing. The second speaker and the second microphone are located on the second side of the electronic device, and the sound propagation path is longer, passing through more reflections and scatterings, and it is easier to generate echoes. In contrast, the first speaker and the first microphone are relatively close, the sound propagation path is short, and the possibility and intensity of echo generation are relatively small. Therefore, when processing the second audio data, a larger echo processing parameter needs to be set to more effectively suppress echoes.
[0143] The spatial environment where the electronic device is located will affect the generation of echoes. Due to the positional relationship between the second speaker and the second microphone, the sound propagation process may be reflected by more obstacles, resulting in more complex and obvious echoes. Therefore, a larger echo processing parameter is required to cope with this situation to ensure that the echoes in the processed audio are effectively suppressed.
[0144] Step 170: Perform operations such as gain adjustment and mixing on the in-ear monitor data, the processed second audio data, and the background audio data, so that they match each other in terms of volume, tone, etc., and finally merge into a complete target audio data.
[0145] Fusion processing: Integrate multiple audio data from different sources or processed in different ways to generate a new and complete audio data. For example, combine the in-ear monitor data, the processed second audio data, and the background audio data to form the final target audio data.
[0146] By performing fusion processing on multiple audio data, the obtained target audio data can comprehensively integrate sound information from various aspects, including both clear user singing voices and appropriate background sounds, and can be used for subsequent audio storage, sharing, or live streaming, etc., providing high-quality KTV works for users.
[0147] In a possible embodiment, step 170 may specifically include the following steps:
[0148] Obtain vocal data based on the in-ear monitor data and the processed second audio data;
[0149] Generate target audio data based on the background audio data and the vocal data.
[0150] Vocal data: Audio data mainly containing user voice information obtained by processing and integrating in-ear monitor data and processed second audio data. It is relatively pure user voice content after removing various interferences such as howling, echo, and other vocal interferences.
[0151] Target audio data: The final generated complete audio data containing background audio data and vocal data, which can be used for subsequent operations such as playback, saving, and sharing. It is the output result of the entire audio processing process and can meet the needs of users in scenarios such as KTV singing and live streaming.
[0152] The in-ear monitor data is obtained by performing a series of processes such as noise reduction, howling processing, and echo suppression on the first audio data collected by the first microphone, and mainly contains the sound information collected by the user through the first microphone. The processed second audio data is obtained by performing similar noise reduction, howling processing, and echo suppression on the second audio data collected by the second microphone, and also contains the user's sound information and some background sound information.
[0153] Analyze and process the in-ear monitor data and the processed second audio data, such as using techniques like audio mixing and duplicate removal. It may match and fuse the corresponding time segments in the two audio data according to the audio timeline, remove the duplicate parts, and further remove the residual interference signals, thereby obtaining relatively pure vocal data. For example, by calculating the amplitude and phase differences between the two audio data at different time points to determine which parts are duplicate and which parts are the user's sound information to be retained.
[0154] Exemplarily, after KTV singing, the audio data processed by the first microphone and the second microphone is input into a dedicated AI processing module to provide a data basis for subsequent in-depth processing. At this time, since the KTV singing has ended, there is no need to consider the latency and power consumption issues of real-time processing, and more complex and powerful processing capabilities can be utilized.
[0155] The large AI model has powerful data analysis capabilities and can perform detailed comparison and analysis on the data of the first microphone and the second microphone channels. By analyzing the differences in amplitude, frequency, phase, etc. between the two, the characteristic distributions of the accompaniment music and the vocals are found, providing a basis for subsequent processing.
[0156] Based on the results of data difference analysis, the AI large model can more accurately identify the accompaniment music component in the signal. Using its deep learning algorithm and a large amount of audio data training experience, adopting more advanced filtering, separation and other technologies, the accompaniment music is more thoroughly separated and eliminated from the audio signal, leaving only a purer vocal component.
[0157] During the echo processing at the front end, some detailed vocal features may be lost due to various reasons, such as tail sounds, small signals, etc. Through learning and understanding a large amount of vocal data, the AI large model can predict and compensate for these lost vocal features according to the characteristics of the current audio signal. By adjusting parameters such as the amplitude, frequency, and phase of the audio, the lost features are re-added to the audio, making the vocals in the audio more complete and natural, and finally updating the target audio data to obtain a cleaner and higher-quality audio output.
[0158] Background audio data is usually prepared in advance, such as the accompaniment music during karaoke singing, etc., which provides a basic background environment for the audio. Vocal data is the voice information of the user. The system will mix the vocal data and the background audio data, and by adjusting parameters such as the volume ratio and balance of the two, make them blend with each other to form a harmonious whole. For example, in the karaoke scene, the volume of the accompaniment music and the user's singing will be appropriately adjusted according to the user's needs and the characteristics of the audio, so that the two match perfectly, and finally the target audio data is generated.
[0159] By obtaining vocal data based on the in-ear monitor data and the processed second audio data, the user voice information from different microphones can be integrated, removing the interference and duplicate parts, and obtaining purer and more complete vocal data. This makes the user's singing clearer and more prominent, improves the quality of the user's voice in the audio, and provides high-quality material for the subsequent fusion with the background audio.
[0160] The generated target audio data can perfectly combine the background audio and the user's vocals, providing an audio work with a good auditory experience. In the karaoke scene, the user can obtain an audio with both appropriate accompaniment and clear vocals of their own, meeting the user's needs for recording, sharing, etc. At the same time, this fusion processing also makes the audio more coordinated and natural as a whole, improving the quality and usability of the audio.
[0161] In an embodiment of the present application, the first microphone is used to collect user audio data and background audio data to obtain first audio data, providing a complete data basis for subsequent audio processing, so as to accurately feedback the user's voice to the user. The first audio data is processed to obtain in-ear monitor data, which can remove the noise components in the first audio data and improve the audio quality. The in-ear monitor data is played through the first speaker. Since the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link, the audio transmission delay is reduced through the hardware link, enabling the user to hear their own voice better aligned with the background audio data and meeting the user's demand for real-time in-ear monitoring.
[0162] An embodiment of the present application also provides an audio playback system, which includes:
[0163] A first microphone, configured to collect user audio data and background audio data to obtain first audio data;
[0164] A real-time in-ear monitor module, configured to process the first audio data to obtain in-ear monitor data;
[0165] A first speaker, configured to play the in-ear monitor data;
[0166] Wherein, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
[0167] The real-time in-ear monitor module processes the first audio data to obtain in-ear monitor data. The processing process may include operations such as noise reduction and audio enhancement, enabling the user to hear their optimized own voice in real time through the first speaker, achieving real-time feedback and facilitating the user to adjust their voice in a timely manner. The in-ear monitor data is looped back to the first speaker through a hardware link, ensuring low latency and stable audio transmission, and enabling the user to obtain a good real-time auditory experience.
[0168] Through the real-time in-ear monitor module, the user can hear their processed own voice in real time. In scenarios such as singing, dubbing, and public speaking, they can better grasp their vocal state, adjust the pitch, rhythm, intonation, etc. in a timely manner, improve the performance or expression effect, and thus enhance the user's experience.
[0169] In a possible embodiment, the system further includes:
[0170] A second speaker, configured to play background audio data;
[0171] A second microphone, configured to collect the user audio data and the background audio data to obtain second audio data;
[0172] An application module is used to perform noise reduction processing on the second audio data to obtain the processed second audio data; and, to perform fusion processing on the in-ear monitor data, the processed second audio data, and the background audio data to obtain target audio data.
[0173] The first speaker and the first microphone are disposed on the first side of the electronic device, and the second speaker and the second microphone are disposed on the second side of the electronic device.
[0174] The first microphone and the second microphone respectively collect user audio data and background audio data. The first microphone collects and forms first audio data for real-time in-ear monitoring; the second microphone collects and forms second audio data for subsequent processing by the application module. Since the two microphones are disposed on different sides of the electronic device, sounds can be collected from different positions, which helps to obtain audio information more comprehensively and accurately.
[0175] The application module performs noise reduction processing on the second audio data to further remove background noise and improve the audio quality. Then, the in-ear monitor data, the processed second audio data, and the background audio data are subjected to fusion processing, and according to different scenarios and requirements, each audio component is reasonably mixed to form target audio data to meet the needs of users in different audio application scenarios.
[0176] By separately collecting and processing different audio data and performing fusion, the various components of the audio can be flexibly adjusted according to different application scenarios and user requirements. For example, in a karaoke scene, the user's singing voice can be enhanced while retaining an appropriate background music; in a dubbing scene, the user's dubbing voice can be highlighted and unnecessary noise can be removed to make the audio effect more in line with the scene requirements.
[0177] The first microphone and the first speaker are disposed on the first side of the electronic device, and the second microphone and the second speaker are disposed on the second side. This layout can utilize the spatial position difference to collect audio data more comprehensively, and may produce a certain spatial sound effect during audio playback, enhancing the three-dimensional sense and immersion of the audio, and bringing a better auditory experience to the user.
[0178] Performing noise reduction processing on the second audio data and noise reduction operations on the first audio data during in-ear monitor processing can effectively remove background noise, improve the clarity and purity of the audio, and make the quality of the finally output target audio data higher, enabling both the user's voice and the background audio to be presented more clearly.
[0179] The following combines Figure 3 to describe the electronic device provided in the embodiments of the present application:
[0180] Application layer: This is where various application programs directly used by users are located, such as KTV applications, recording applications, etc. It is responsible for initiating audio-related operation requests and receiving the processed audio data for functions such as playback and display.
[0181] Framework layer: Parses and schedules the requests from the application layer, manages and coordinates functional modules such as the human voice recording path, accompaniment recording path, and real-time ear monitor, ensuring the orderly progress of the audio processing flow.
[0182] Audio Digital Signal Processor (ADSP) layer: Mainly performs digital processing of audio signals, including specific processing functions such as echo cancellation, howling suppression, external speaker sound effect processing, and external speaker protection. It is the core link of audio signal processing.
[0183] Peripheral layer: The bottom layer, which is the hardware devices for actually collecting and playing audio, namely the first microphone, the second microphone, the first speaker, and the second speaker, responsible for the collection input and playback output of audio signals.
[0184] The following explains each path and functional module:
[0185] Regarding the human voice recording path: The first microphone and the second microphone collect human voice audio data, and these raw data are accompanied by interference such as ambient sound and echo in the surrounding environment. The collected data first enters the howling suppression module. By analyzing features such as the frequency of the audio signal, it detects and suppresses possible howling signals to avoid the appearance of sharp and ear-piercing sounds.
[0186] Then it enters the echo cancellation module. Using signal processing algorithms, it removes the echo generated due to sound reflection, etc., making the collected human voice purer. The processed human voice data is uploaded to the framework layer for use by the application layer.
[0187] Regarding the accompaniment recording path: It mainly involves obtaining accompaniment audio data from the device, passing through the external speaker sound effect module to perform sound effect enhancement and other processing on the accompaniment audio, such as adjusting volume balance, timbre, etc. Then, through the external speaker protection module, it prevents damage to the speaker due to excessive volume, etc. After that, it is transmitted to the framework layer and can be used for subsequent fusion operations with the processed human voice data.
[0188] Regarding the real-time ear monitor module: The real-time ear monitor function relies on the human voice data collected and processed by the human voice recording path, quickly transmits the processed human voice data back to the first speaker or the second speaker to achieve real-time ear monitoring. It enables users to hear their own voices in a timely manner and adjust their singing or expression. The key lies in low-latency processing to ensure sound synchronization.
[0189] The following explains each data flow:
[0190] Recording data stream: Formed by collecting audio signals such as human voices by the first microphone and the second microphone, and uploaded after being processed by howling suppression, echo cancellation, etc. It is the transmission path from the original human voice signal to the processed signal.
[0191] Playback data stream: The processed accompaniment audio data in the accompaniment recording path, and the complete audio data after being fused with the processed human voice data, are transmitted through this path to the speaker for playback.
[0192] In-ear monitor data stream: The human voice data processed in the human voice recording path is transmitted to the speaker through the real-time in-ear monitor module to achieve the real-time in-ear monitor function.
[0193] Control flow: Transmits control instructions between layers. For example, instructions such as audio acquisition and processing parameter settings in the application layer are transmitted through the control flow to the framework layer and the ADSP layer to achieve control of the entire audio processing process.
[0194] In the embodiment of the present application, the execution subject of the provided audio playback method can be an audio playback device. In the embodiment of the present application, taking the audio playback device executing the audio playback method as an example, the audio playback device provided by the embodiment of the present application is described.
[0195] Figure 4 It is a block diagram of an audio playback device provided by an embodiment of the present application. The device 400 includes:
[0196] The acquisition module 410 is configured to collect user audio data and background audio data through the first microphone to obtain first audio data;
[0197] The processing module 420 is configured to process the first audio data to obtain in-ear monitor data;
[0198] The playback module 430 is configured to play the in-ear monitor data through the first speaker;
[0199] Among them, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
[0200] In a possible embodiment, the playback module 430 is further configured to play the background audio data through the second speaker;
[0201] The acquisition module 410 is further configured to collect the user audio data and the background audio data through the second microphone to obtain second audio data;
[0202] The processing module 420 is further configured to perform noise reduction processing on the first audio data to obtain the in-ear monitor data;
[0203] The processing module 420 is further configured to perform noise reduction processing on the second audio data to obtain the processed second audio data;
[0204] The apparatus 400 may further include:
[0205] A fusion module, configured to perform fusion processing on the in-ear monitor data, the processed second audio data, and the background audio data to obtain target audio data.
[0206] In a possible embodiment, the processing module 420 is specifically configured to:
[0207] Perform howling suppression on the first audio data according to a first howling processing parameter, a first link delay, and a user voice feature to obtain the first audio data after howling suppression; the first link delay is used to describe the link delay of the user audio data from being played by the first speaker to being collected by the first microphone;
[0208] Perform echo suppression processing on the first audio data after howling suppression according to a first echo processing parameter to obtain in-ear monitor data.
[0209] In a possible embodiment, the processing module 420 is specifically configured to:
[0210] Perform howling suppression on the second audio data according to a second howling processing parameter, a second link delay, and a user voice feature to obtain the second audio data after howling suppression; the second link delay is used to describe the link delay of the user audio data from being played by the second speaker to being collected by the second microphone;
[0211] Perform echo suppression processing on the second audio data after howling suppression according to a second echo processing parameter to obtain the processed second audio data.
[0212] In a possible embodiment, the first speaker and the first microphone are disposed on a first side of the electronic device, and the second speaker and the second microphone are disposed on a second side of the electronic device;
[0213] The second howling processing parameter is less than the first howling processing parameter corresponding to the first audio data;
[0214] The second echo processing parameter is greater than the first echo processing parameter corresponding to the first audio data.
[0215] In a possible embodiment, the fusion module is specifically configured to:
[0216] Obtain human voice data according to the in-ear monitor data and the processed second audio data;
[0217] Generate target audio data according to the background audio data and the human voice data.
[0218] In an embodiment of the present application, the first microphone is used to collect user audio data and background audio data to obtain first audio data, providing a complete data basis for subsequent audio processing, so as to accurately feedback the user's voice to the user. The first audio data is processed to obtain in-ear monitor data, which can remove the noise components in the first audio data and improve the audio quality. The in-ear monitor data is played through the first speaker. Since the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link, the audio transmission delay is reduced through the hardware link, enabling the user to hear their own voice better aligned with the background audio data and meeting the user's requirement for real-time in-ear monitoring.
[0219] The audio playback device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application does not make specific limitations.
[0220] The audio playback device in the embodiment of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiment of the present application does not make specific limitations.
[0221] The audio playback device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, it will not be elaborated here.
[0222] Optionally, as Figure 5As shown in the figure, an embodiment of the present application further provides an electronic device 510, including a processor 511, a memory 512, a program or instruction stored on the memory 512 and executable on the processor 511. When the program or instruction is executed by the processor 511, it implements the steps of any of the above audio playback method embodiments and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0223] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0224] Figure 6 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.
[0225] The electronic device 600 includes, but is not limited to: a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, and a processor 610, etc.
[0226] Those skilled in the art can understand that the electronic device 600 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 610 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 6 The structure of the electronic device shown in [the figure] does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated herein.
[0227] Among them, the input unit 604 is used to collect user audio data and background audio data through a first microphone to obtain first audio data;
[0228] The processor 610 is used to process the first audio data to obtain in-ear monitor data;
[0229] The audio output unit 603 is used to play the in-ear monitor data through a first speaker;
[0230] Among them, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
[0231] Optionally, the audio output unit 603 is further used to play background audio data through a second speaker;
[0232] The input unit 604 is further used to collect the user audio data and the background audio data through a second microphone to obtain second audio data;
[0233] The processor 610 is further configured to perform noise reduction processing on the first audio data to obtain the in-ear monitor data;
[0234] The processor 610 is further configured to perform noise reduction processing on the second audio data to obtain the processed second audio data;
[0235] The processor 610 is further configured to perform fusion processing on the in-ear monitor data, the processed second audio data, and the background audio data to obtain the target audio data.
[0236] Optionally, the processor 610 is further configured to perform howling suppression on the first audio data according to the first howling processing parameter, the first link delay, and the user voice feature to obtain the first audio data after howling suppression; the first link delay is used to describe the link delay of the user audio data from being played by the first speaker to being collected by the first microphone;
[0237] The processor 610 is further configured to perform echo suppression processing on the first audio data after howling suppression according to the first echo processing parameter to obtain the in-ear monitor data.
[0238] Optionally, the processor 610 is further configured to perform howling suppression on the second audio data according to the second howling processing parameter, the second link delay, and the user voice feature to obtain the second audio data after howling suppression; the second link delay is used to describe the link delay of the user audio data from being played by the second speaker to being collected by the second microphone;
[0239] The processor 610 is further configured to perform echo suppression processing on the second audio data after howling suppression according to the second echo processing parameter to obtain the processed second audio data.
[0240] Optionally, the first speaker and the first microphone are disposed on a first side of the electronic device, and the second speaker and the second microphone are disposed on a second side of the electronic device;
[0241] The second howling processing parameter is less than the first howling processing parameter corresponding to the first audio data;
[0242] The second echo processing parameter is greater than the first echo processing parameter corresponding to the first audio data.
[0243] In a possible embodiment, the processor 610 is further configured to obtain human voice data according to the in-ear monitor data and the processed second audio data;
[0244] The processor 610 is further configured to generate target audio data according to the background audio data and the human voice data.
[0245] In an embodiment of the present application, the first microphone is used to collect user audio data and background audio data to obtain first audio data, providing a complete data basis for subsequent audio processing, so as to accurately feedback the user's voice to the user. The first audio data is processed to obtain in-ear monitor data, which can remove the noise components in the first audio data and improve the audio quality. The in-ear monitor data is played through the first speaker. Since the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link, the audio transmission delay is reduced through the hardware link, enabling the user to hear their own voice better aligned with the background audio data and meeting the user's demand for real-time in-ear monitoring.
[0246] It should be understood that in an embodiment of the present application, the input unit 604 may include a Graphics Processing Unit (GPU) 6041 and a microphone 6042. The graphics processor 6041 processes the image data of static pictures or video images obtained by an image capture device (such as a camera) in a video image capture mode or an image capture mode. The display unit 606 may include a display panel 6061, and the display panel 6061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also referred to as a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. The other input devices 6072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here. The memory 609 can be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 610 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, and the modem processor mainly processes wireless communications. It can be understood that the above modem processor may not be integrated into the processor 610.
[0247] The memory 609 can be used to store software programs and various data. The memory 609 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 609 can include volatile memory or non-volatile memory, or the memory 609 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 609 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memory.
[0248] The processor 610 may include one or more processing units; optionally, the processor 610 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 610 either.
[0249] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above audio playback method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0250] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.
[0251] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above audio playback method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0252] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0253] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above audio playback method embodiment, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0254] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described method may be executed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0255] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0256] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. An audio playback method, characterized in that, The method includes: Collecting user audio data and background audio data through a first microphone to obtain first audio data; Processing the first audio data to obtain in-ear monitor data; Playing the in-ear monitor data through a first speaker; Wherein, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
2. The method according to claim 1, wherein The method further includes: Playing the background audio data through a second speaker; Collecting the user audio data and the background audio data through a second microphone to obtain second audio data; Performing noise reduction processing on the first audio data to obtain the in-ear monitor data; Performing noise reduction processing on the second audio data to obtain processed second audio data; Performing fusion processing on the in-ear monitor data, the processed second audio data, and the background audio data to obtain target audio data.
3. The method according to claim 2, wherein The performing noise reduction processing on the first audio data to obtain the in-ear monitor data includes: Performing howling suppression on the first audio data according to a first howling processing parameter, a first link delay, and user voice characteristics to obtain the first audio data after howling suppression; the first link delay is used to describe the link delay of the user audio data from being played by the first speaker to being collected by the first microphone; Performing echo suppression processing on the first audio data after howling suppression according to a first echo processing parameter to obtain in-ear monitor data.
4. The method according to claim 2, wherein The performing noise reduction processing on the second audio data to obtain the processed second audio data includes: Performing howling suppression on the second audio data according to a second howling processing parameter, a second link delay, and user voice characteristics to obtain the second audio data after howling suppression; the second link delay is used to describe the link delay of the user audio data from being played by the second speaker to being collected by the second microphone; Performing echo suppression processing on the second audio data after howling suppression according to a second echo processing parameter to obtain the processed second audio data.
5. The method according to claim 4, wherein The first speaker and the first microphone are disposed on a first side of the electronic device, and the second speaker and the second microphone are disposed on a second side of the electronic device; The second howling processing parameter is less than the first howling processing parameter corresponding to the first audio data; The second echo processing parameter is greater than the first echo processing parameter corresponding to the first audio data.
6. The method according to claim 2, wherein The performing fusion processing on the in-ear monitor data, the processed second audio data, and the background audio data to obtain target audio data includes: Obtaining human voice data according to the in-ear monitor data and the processed second audio data; Generating target audio data according to the background audio data and the human voice data.
7. An audio playback system, characterized in that, The system includes: A first microphone, configured to collect user audio data and background audio data to obtain first audio data; A real-time in-ear monitor module, configured to process the first audio data to obtain in-ear monitor data; A first speaker, configured to play the in-ear monitor data; Wherein, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
8. The system according to claim 7, wherein The system further includes: A second speaker for playing background audio data; A second microphone for collecting the user audio data and the background audio data to obtain second audio data; An application module for performing noise reduction processing on the second audio data to obtain processed second audio data; and for performing fusion processing on the in-ear monitor data, the processed second audio data, and the background audio data to obtain target audio data; The first speaker and the first microphone are disposed on a first side of the electronic device, and the second speaker and the second microphone are disposed on a second side of the electronic device.
9. An audio playback device, characterized in that, The apparatus includes: A collection module for collecting user audio data and background audio data through a first microphone to obtain first audio data; A processing module for processing the first audio data to obtain in-ear monitor data; A playback module for playing the in-ear monitor data through a first speaker; Wherein, the in-ear monitor data is looped back from the first microphone to the first speaker through a hardware link.
10. An electronic device, characterized in that, Comprising a processor and a memory, the memory stores a program or instruction executable on the processor, and when the program or instruction is executed by the processor, the steps of the audio playback method according to any one of claims 1-6 are implemented.
11. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the audio playback method according to any one of claims 1-6 are implemented.