Audio data processing method, processing system and processing device
By adjusting the time difference between the audio playback retrieval function and the microphone acquisition function, the audio data phase is aligned and combined, thus solving the problem of audio data beat discrepancies in smart devices and improving the accuracy of speech recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-31
- Publication Date
- 2026-03-27
AI Technical Summary
In speech recognition, the audio data collected by smart devices contains noise in addition to the target audio data, leading to inaccurate recognition and misrecognition. Furthermore, due to excessive CPU load, there are beat differences in the audio data, resulting in phase misalignment and severe jitter.
By determining the time difference between the audio playback recapture function and the microphone acquisition function, the audio data is adjusted to align their phases, and the aligned audio data is combined for echo cancellation.
Phase alignment of audio data was achieved, eliminating beat problems and improving the accuracy and stability of speech recognition.
Smart Images

Figure CN116564327B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio processing, in particular to an audio data processing method, a processing system and a processing device. BACKGROUND
[0002] In speech recognition, the audio data collected by the intelligent device often contains noise in addition to the target audio data, such as the audio data (reference audio data) played by the audio playback module of the intelligent device. At this time, if the collected audio data is subjected to speech recognition, there will be problems such as inaccurate recognition, misrecognition, and failure to recognize, and due to the influence of the system or software in the intelligent device on the short-time CPU occupancy rate, there is a beat problem between the collected audio data and the reference audio data, wherein the beat phenomenon can be understood as the collected audio data and the reference audio data being out of phase, and since the CPU load is variable, the severity of the beat of the data collected each time is different, that is, the size of the beat has jitter. SUMMARY
[0003] The present application provides an audio data processing method, a processing system and a processing device, which can solve the problem of beat and phase misalignment between two audio data, especially the problem of misalignment with jitter caused by excessive CPU load, and thus can use the combined audio data obtained by combination for echo cancellation processing.
[0004] To solve the above technical problems, one technical solution adopted by the present application is to provide an audio data processing method, which comprises: determining the time difference between starting the audio playback back collection function and starting the microphone collection function by using the audio driver, and obtaining the first audio data collected by the audio playback back collection function and the second audio data collected by the microphone collection function; adjusting one of the first audio data and the second audio data based on the time difference, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned; combining the unadjusted audio data and the adjusted audio data to obtain combined audio data; and performing echo cancellation based on the combined audio data.
[0005] Among them, adjusting one of the first audio data and the second audio data based on the time difference, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned, comprises: based on the time difference, determining one of the first audio data and the second audio data as the audio data to be adjusted, and the other as the reference audio data; adjusting the audio data to be adjusted so that the adjusted audio data to be adjusted is phase-aligned with the reference audio data.
[0006] The time difference is a difference between a time of starting the audio playback collection function and a time of starting the microphone collection function.
[0007] The time difference is a difference between a time of starting the audio playback collection function and a time of starting the microphone collection function.
[0008] The time difference is a difference between a time of starting the audio playback collection function and a time of starting the microphone collection function.
[0009] The time difference is a difference between a time of starting the audio playback collection function and a time of starting the microphone collection function.
[0010] The time difference is a difference between a time of starting the audio playback collection function and a time of starting the microphone collection function.
[0011] The time difference is a difference between a time of starting the audio playback collection function and a time of starting the microphone collection function.
[0012] The time difference is a difference between a time of starting the audio playback collection function and a time of starting the microphone collection function.
[0013] The adjusting one of the first audio data and the second audio data based on the time difference comprises: obtaining an adjustment parameter received by an adjustment interface; and adjusting the time difference by using the adjustment parameter.
[0014] To solve the above technical problems, another technical solution adopted by the present application is to provide an audio data processing system, which comprises an audio playback and collection module, a microphone collection module, an audio driver and a processing module.
[0015] The audio playback and collection module is configured to collect the first audio data; the microphone collection module is configured to collect the second audio data; the audio driver is connected to the audio playback and collection module and the microphone collection module, and is configured to determine a time difference between an audio playback and collection function of the audio playback and collection module and a microphone collection function of the microphone collection module, and to obtain the first audio data collected by the audio playback and collection module and the second audio data collected by the microphone collection module; the processing module is connected to the audio playback and collection module, the microphone collection module and the audio driver, and is configured to adjust one of the first audio data and the second audio data based on the time difference, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned; and to combine the unadjusted audio data and the adjusted audio data to obtain combined audio data; and to perform echo cancellation based on the combined audio data.
[0016] The processing system further comprises a first buffer, a second buffer and a third buffer. The first buffer is connected to the audio playback and collection module, the audio driver and the processing module, and is configured to buffer the first audio data; the second buffer is connected to the microphone collection module, the audio driver and the processing module, and is configured to buffer the second audio data.
[0017] The processing module is configured to read the first audio data from the first buffer and the second audio data from the second buffer, and to adjust one of the first audio data and the second audio data based on the time difference, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned; and to combine the unadjusted audio data and the adjusted audio data to obtain combined audio data, and to buffer the combined audio data in the third buffer, which is connected to the processing module.
[0018] To solve the above technical problems, another technical solution adopted by the present application is to provide an audio data processing device, which comprises a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program to implement the above-mentioned audio data processing method.
[0019] The beneficial effects of the present application are: different from the prior art, the audio data processing method provided by the present application can determine the time difference between starting the audio playback back sampling function and starting the microphone collection function, and obtain the first audio data collected by the audio playback back sampling function and the second audio data collected by the microphone collection function; then, based on the time difference, one of the first audio data and the second audio data is adjusted, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned; the unadjusted audio data and the adjusted audio data are combined to obtain combined audio data; finally, echo cancellation is performed based on the combined audio data. In the above manner, the first audio data and the second audio data can be aligned, the problem of beat and phase misalignment between the audio data collected by the audio playback back sampling function and the audio data collected by the microphone collection function can be solved, and the two audio data are further combined to obtain combined audio data which can be used for echo cancellation to eliminate the corresponding first audio data in the second audio data. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:
[0021] Figure 1 is a flowchart of the first embodiment of the audio data processing method provided by the present application;
[0022] Figure 2 is a flowchart of an embodiment of step 12 provided by the present application;
[0023] Figure 3 is a flowchart of an embodiment of step 22 provided by the present application;
[0024] Figure 4 is a schematic diagram of an embodiment of the audio data combination method provided by the present application;
[0025] Figure 5 is a flowchart of the second embodiment of the audio data processing method provided by the present application;
[0026] Figure 6 is a structural schematic diagram of the first embodiment of the audio data processing system provided by the present application;
[0027] Figure 7is a structural schematic diagram of a second embodiment of an audio data processing system provided by the present application;
[0028] Figure 8 is a structural schematic diagram of an embodiment of an audio data processing device provided by the present application;
[0029] Figure 9 is a structural schematic diagram of an embodiment of a computer readable storage medium provided by the present application. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0031] In voice recognition, the audio data collected by the intelligent device often contains noise in addition to the target audio data, such as the audio data (reference audio data) played by the audio playback module of the intelligent device. At this time, if the collected audio data is subjected to voice recognition, there will be problems such as inaccurate recognition, misrecognition, and failure to recognize, and due to the influence of the system or software in the intelligent device on the short-time CPU occupancy rate, there is a beat problem between the collected audio data and the reference audio data, wherein the beat phenomenon can be understood as that the collected audio data and the reference audio data are not aligned in phase, and since the CPU load is variable, the severity of the beat of the data collected each time is different, that is, the size of the beat is dithered. Based on this, the present application proposes the following any embodiment to solve the at least one technical problem.
[0032] Referring to Figure 1 , Figure 1 is a flow schematic diagram of a first embodiment of an audio data processing method provided by the present application, and the method comprises the following steps:
[0033] Step 11: determining a time difference between starting the audio playback back sampling function and starting the microphone collection function by using the audio driver, and obtaining first audio data collected by the audio playback back sampling function and second audio data collected by the microphone collection function.
[0034] Specifically, the time difference is the difference between the time of starting the audio playback back sampling function of the audio playback back sampling module and the time of starting the microphone collection function of the microphone collection module.
[0035] The audio playback back sampling module is a module for collecting audio data of the audio playback module. The audio playback module is referred to as module A, and the audio playback back sampling module is referred to as module B. Module A can be started at any time. If module B is started when module A has been started, the audio data collected by module B is downlink audio data. If module B is started when module A has not been started, the audio data collected by module B is silent data. For example, the audio data played by module A is "Xiaoming Xiaoming, where are you?" If module B is started when module A has been started, the audio data collected by module B can be "Ming Xiaoming, where are you?", "Xiaoming, where are you?", or other data. If module B is started when module A has not been started, the audio data collected by module B is silent data.
[0036] The microphone collection module is used to collect audio data output by a user or other audio devices / modules. It is worth noting that the second audio data collected by the microphone collection module can include the first audio data. For example, the first audio data collected by the audio playback back sampling module is "where are you?", and at the same time, the user inputs the voice "Xiaoming Xiaoming". The second audio data collected by the microphone collection module is the audio data of "Xiaoming Xiaoming" and "where are you?" mixed together.
[0037] Optionally, the audio playback back sampling module can be an auxiliary module in a vehicle-mounted device, a smart phone, a notebook computer, or a smart wearable device. The microphone collection module can be a MIC (Microphone) device of a sound card in the smart phone, the notebook computer, or the vehicle-mounted device. In other words, the audio playback back sampling function can be a function of collecting audio playback data, which is realized by starting / turning on an auxiliary module in the vehicle-mounted device, the smart phone, the notebook computer, or the smart wearable device. The microphone collection function can be an audio data collection function, which is realized by starting / turning on a MIC device of a sound card.
[0038] In some embodiments, the audio playback back sampling function and the microphone collection function are started by a software program / programming code.
[0039] In some embodiments, the audio driver is an Advanced-Linux-Sound-Architecture System-on-Chip driver (ALSA SoC driver). The Advanced-Linux-Sound-Architecture System-on-Chip driver can record the starting time of the audio playback back sampling function and the starting time of the microphone collection function to determine the time difference. For example, the starting time of the audio playback back sampling function is 18:00, and the starting time of the microphone collection function is 18:02. At this time, the time difference is 2 minutes.
[0040] In some embodiments, the time when the audio playback back sampling function is enabled and the time when the microphone sampling function is enabled can be ensured to be accurate by setting a spin lock (spin_lock_irqsave).
[0041] After the audio playback back sampling function and the microphone sampling function are enabled, the audio playback back sampling function and the microphone sampling function will each collect corresponding audio data. For example, the first audio data collected by the audio playback back sampling function and the second audio data collected by the microphone sampling function. Among the second audio data, there are usually other audio data in addition to the audio data of the audio playback. Therefore, echo cancellation is needed.
[0042] Step 12: Adjust one of the first audio data and the second audio data based on the time difference, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned.
[0043] It can be understood that due to the time difference, the first audio data and the second audio data may not be phase-aligned. It can be understood that the audio data is actually propagated in the form of sound waves, and the so-called phase refers to the position of a sound wave at different cycle stages, which is represented by 0 degrees, 90 degrees, 180 degrees, 270 degrees, and 360 degrees. The problem of phase comes from the fact that more than two collection modules collect audio data from the same sound source. Because of the difference between the collection modules, a single sound source will have a time difference for different collection modules. The audio data that should have been received at the same time point appears at different time points in different collection modules. This time difference causes a phase difference, i.e., a phase misalignment.
[0044] Therefore, it is only necessary to adjust one of the first audio data and the second audio data to achieve phase alignment with the other.
[0045] In some embodiments, referring to Figure 2 , step 12 can be the following flow:
[0046] Step 21: Based on the time difference, one of the first audio data and the second audio data is determined as the audio data to be adjusted, and the other is determined as the reference audio data.
[0047] It can be understood that the time difference is the difference between the time when the audio playback back sampling function is enabled and the time when the microphone sampling function is enabled.
[0048] In some embodiments, one of the first audio data and the second audio data is determined as the reference audio data based on the difference between the time when the audio playback back sampling function is enabled and the time when the microphone sampling function is enabled, and the other audio data is adjusted based on the reference audio data.
[0049] As the difference between the time when the audio playback back-sampling function is started and the time when the microphone collection function is started is negative, the starting time of the audio playback back-sampling function is prior to the microphone collection function. At this time, the second audio data collected by the microphone collection function will lose the part of the audio data played at the beginning collected by the audio playback back-sampling function. Therefore, the first audio data collected by the audio playback back-sampling function is taken as the reference audio data, and the second audio data collected by the microphone collection function is taken as the audio data to be adjusted.
[0050] That is, the audio data obtained by the function started first is taken as the reference to adjust the audio data obtained by the function started later.
[0051] In some embodiments, when the difference between the time when the audio playback back-sampling function is started and the time when the microphone collection function is started is non-negative (zero or positive), it indicates that the time when the second audio data collected by the microphone collection function after the microphone collection function is started will be slightly later than the time when the first audio data collected by the audio playback back-sampling function after the audio playback back-sampling function is started.
[0052] When the difference is zero, it indicates that the audio playback back-sampling function and the microphone collection function are started at the same time. It is worth noting that due to the influence of air propagation and other factors, the time when the first audio data collected by the audio playback back-sampling function will be prior to the time when the second audio data collected by the microphone collection function. At this time, the first audio data and the second audio data are not synchronized in time, so the second audio data is taken as the reference audio data, and the first audio data is taken as the audio data to be adjusted. Of course, in other embodiments, the difference in the installation position of the hardware device can also cause the phase misalignment of the first audio data and the second audio data.
[0053] In some embodiments, before adjusting one of the first audio data and the second audio data, the adjustment parameters received by the adjustment interface can be obtained, and the time difference determined above is adjusted by using the adjustment parameters, and then one of the first audio data and the second audio data is adjusted based on the adjusted time difference, so that the phase alignment between the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data. The adjustment parameters can be manually input, that is, the user can customize the adjustment of the time error caused by the difference in the installation position of the hardware device through the adjustment interface.
[0054] Similarly, when the difference is positive, it means that the microphone collection function is prior to the audio playback collection function, and in addition to the influence of factors such as air propagation, the first audio data and the second audio data are also out of synchronization in time. In order to align the first audio data and the second audio data, the second audio data is taken as the reference audio data, and the first audio data is taken as the audio data to be adjusted. The first audio data is adjusted based on the time difference, so that the first audio data and the second audio data are phase-aligned.
[0055] Step 22: adjusting the audio data to be adjusted so that the adjusted audio data to be adjusted and the reference audio data are phase-aligned.
[0056] It can be understood that phase is a reflection of time characteristics, and phase represents a periodic motion waveform, such as a sine wave in a period, which can change from 0 degrees to 360 degrees. When two audio data are mixed, if the phases are not aligned, phase cancellation may occur, which further leads to signal loss.
[0057] Therefore, when it is determined that the phases of two audio data are not aligned, the two audio data need to be adjusted to achieve phase alignment to avoid signal loss and the like.
[0058] In some embodiments, referring to Figure 3 , step 22 can be the following flow:
[0059] Step 31: determining the number of bytes of difference between the audio data to be adjusted and the reference audio data.
[0060] Specifically, due to the influence of system or software leading to CPU loading fluctuation, there is a difference between the time when the audio playback collection function is started and the time when the microphone collection function is started, which causes the first audio data collected by the audio playback collection function and the second audio data collected by the microphone collection function to have a beat.
[0061] At this time, the number of bytes of difference between the first audio data and the second audio data can be calculated by the time difference, combined with the sampling frequency f, the format, and the number of channels. The specific steps are as follows (not shown in the figure):
[0062] S1: calculating the number of audio data stored by the first audio data in the first time, and calculating the number of audio data collected by the second audio data in the second time.
[0063] Specifically, the formula "storage amount = sampling frequency x sampling bit number x channel number x time / 8" is used to calculate the amount of audio data in a unit of time.
[0064] Wherein, the unit of the sampling frequency is Hertz; the number of channels of the mono sound is 1, the number of channels of the stereo sound is 2, and the number of channels of the four sound is 4; the sampling bit of the mono sound and the stereo sound is 16 bits; and the unit of the time is second (s).
[0065] For example, if the sampling frequency of a record is 44KHz, the sampling bit is 16 bits, and the stereo sound (2 channels), the storage amount of the audio data in one minute (60 seconds) is: 44x1000x16x2x60 / 8=10560000 (bytes).
[0066] S2: Compare the amount of audio data stored in the first time by the first audio data and the amount of audio data stored in the second time by the second audio data to determine the byte difference between the first audio data and the second audio data.
[0067] Since there is a difference between the time when the audio playback back sampling function is started and the time when the microphone collection function is started, the difference is taken as the time difference. If the time difference is negative, it means that the time of the first audio data collected by the audio playback back sampling function is earlier than the time of the second audio data collected by the microphone collection function. At this time, the second time can be determined based on the first time, and the second time is obtained by subtracting the time difference from the first time.
[0068] For example, the first time is 5 seconds, and the time difference is 1 second. Taking the sampling frequency of 44KHz, the sampling bit of 16 bits, and the 2 channels as an example, the storage data amount of the second audio data in 4 seconds is 44x1000x16x2x4 / 8=704000 bytes, and the data amount stored in 5 seconds by the first audio data is 44x1000x16x2x5 / 8=880000 bytes. At this time, the byte difference between the first audio data and the second audio data is 880000-704000=176000 bytes.
[0069] Of course, the "time" in the formula can be directly converted into "time difference" for calculation, such as 44x1000x16x2x1 / 8=176000 bytes.
[0070] In addition, in some embodiments, the time difference is required to be less than or equal to 10 milliseconds, and the time difference is a non-negative number, that is, the time of starting the audio playback back sampling function to collect the first audio data lags behind the time of starting the microphone collection function to collect the second audio data.
[0071] Step 32: Perform data padding operation on the to-be-adjusted audio data based on the byte number, so that the adjusted to-be-adjusted audio data is phase-aligned with the reference audio data.
[0072] Specifically, the start position of the audio data to be adjusted is padded with zeros according to the number of bytes, so that the adjusted audio data to be adjusted is aligned with the reference audio data.
[0073] In some embodiments, when the difference between the time when the audio playback and collection function is started and the time when the microphone collection function is started is negative, the first audio data is taken as the reference audio data, and the second audio data is taken as the audio data to be adjusted. At this time, the corresponding silence data is inserted into the buffer corresponding to the microphone collection function. In some embodiments, the silence data is 0.
[0074] For example, the misaligned first audio data and second audio data are:
[0075] MIC 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 REF 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
[0076] In the above table, REF represents the first audio data, and MIC represents the second audio data. The audio data played / captured is represented by "12345…". As can be seen from the above table, because the microphone collection function is started later than the audio playback and collection function, when the first audio data (1 in the table) of the first audio data is collected, the second audio data has not collected the first audio data played by the first audio data. Therefore, when the first audio data is played to the fifth audio data (5 in the table), the second audio data collects the fifth audio data, indicating that the second audio data is later than the first audio data. At this time, the first audio data is taken as the reference, and the number of bytes of the beat between the first audio data and the second audio data is determined to adjust the second audio data. For example, the audio data collected by the second audio data is padded with zeros at the beginning, as shown in the following table.
[0077] The aligned first audio data and second audio data are:
[0078] MIC 0 0 0 0 5 6 7 8 9 10 11 12 13 14 15 REF 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
[0079] In some embodiments, when the difference between the time when the audio playback and collection function is started and the time when the microphone collection function is started is non-negative (zero or positive), the second audio data is taken as the reference audio data, and the first audio data is taken as the audio data to be adjusted.
[0080] For example, the misaligned first audio data and second audio data are:
[0081] MIC 0 0 1 2 3 4 5 6 7 8 9 10 11 12 13 REF 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
[0082] Wherein, REF represents the first audio data, MIC represents the second audio data, and the audio data played / captured is represented by "123456…". As can be seen from the above table, when the first audio data and the second audio data are time-aligned (the difference is zero), there is still the problem of beat between the first audio data and the second audio data caused by air propagation and other influences, and further alignment operation is still needed. As can be seen from the above table, when the first audio data (which is "1" in the table) is captured by the second audio data, the second audio data has not captured the first and second audio data (which are "1" and "2" in the table) played by the first audio data. Therefore, when the first audio data is played to the third audio data (which is "3" in the table) when the first audio data is captured by the second audio data, it indicates that the first audio data is prior to the second audio data in time. Therefore, the number of bytes of beat between the first audio data and the second audio data is determined based on the second audio data as the reference audio data, and the first audio data is adjusted, as shown in the following table.
[0083] The aligned first audio data and second audio data are:
[0084] MIC 0 0 1 2 3 4 5 6 7 8 9 10 11 12 13 REF 0 0 1 2 3 4 5 6 7 8 9 10 11 12 13
[0085] Step 13: combining the unadjusted audio data and the adjusted audio data to obtain combined audio data.
[0086] It can be understood that if the adjusted audio data is the first audio data and the unadjusted audio data is the second audio data, the adjusted first audio data and the unadjusted second audio data can be combined to obtain combined audio data after the first audio data and the second audio data are phase-aligned. If the adjusted audio data is the second audio data and the unadjusted audio data is the first audio data, the adjusted second audio data and the unadjusted first audio data can be combined to obtain combined audio data after the first audio data and the second audio data are phase-aligned.
[0087] In some embodiments, the first audio data and the second audio data after phase alignment are combined based on an Android system. The Android system includes four layers, namely an APP (Application) layer, a Framework layer, a HAL layer (also called a hardware abstraction layer), and a Kernel layer (also called a kernel layer).
[0088] Specifically, the unadjusted audio data can be combined with the adjusted audio data based on the HAL layer, i.e., the first audio data after phase alignment is combined with the second audio data to obtain combined audio data; or the unadjusted audio data can be combined with the adjusted audio data at the Kernel layer, i.e., the first audio data after phase alignment is combined with the second audio data to obtain combined audio data.
[0089] As the first audio data after phase alignment and the second audio data are combined at the HAL layer, the read thread starting action of the audio playback and microphone collection functions needs to be performed at the HAL layer first, the microphone collection function and the audio playback function enter a waiting state after starting the read thread, then the semaphores in the two read threads are continuously triggered at the HAL layer, so that the two read threads start reading at almost the same time, then the logic operation of ALSA (Advanced Linux Sound Architecture) is used to return the read data when the read data reaches a preset number. In addition, even if two read operation instructions are continuously issued, there is still a time difference in the execution of the two instructions, and the time difference changes with the difference in system or software loading, thereby causing different time differences.
[0090] Notably, if the first audio data after phase alignment and the second audio data need to be combined at the HAL layer, the first audio data and the second audio data are read from a sound card device (sound card driver); if the first audio data after phase alignment and the second audio data need to be combined at the Kernel layer, the first audio data and the second audio data are directly read from hardware.
[0091] In some embodiments, a preset combined audio data format is first obtained, and then the first audio data after phase alignment and the second audio data are combined based on the obtained preset combined audio data format to obtain combined audio data.
[0092] The combined audio data format can be to splice double audio zones into 4 channels or to splice four audio zones into 6 channels. Double audio zones represent left and right two audio zones, and four audio zones represent front left, front right, rear left, and rear right four audio zones; 4 channels refer to that one frame of combined data is 4 channels, one frame of data includes two second audio data and two first audio data, and 6 channels refer to that one frame of combined data is 6 channels, one frame of data includes four second audio data and two first audio data.
[0093] In some embodiments, the preset combined audio data format is composed of at least the following bytes, wherein the first byte is for the first left channel data of the second audio data, the second byte is for the first right channel data of the second audio data, the third byte is for the first left channel data of the first audio data, the fourth byte is for the first right channel data of the first audio data, the fifth byte is for the second left channel data of the second audio data, the sixth byte is for the second right channel data of the second audio data, the seventh byte is for the second left channel data of the first audio data, and the eighth byte is for the second right channel data of the first audio data. As shown in Figure 4 Figure 4 is an embodiment of the combined manner of the first audio data and the second audio data.
[0094] The data format of the first audio data (2 channels) is:
[0095] REF_1L REF_1R REF_2L REF_2R
[0096] REF represents the first audio data, and the storage manner of the audio data is LRLR (left-right-left-right).
[0097] The data format of the second audio data (2 channels) is:
[0098] MIC_1L MIC_1R MIC_2L MIC_2R
[0099] MIC represents the second audio data, and the storage manner of the audio data is LRLR (left-right-left-right).
[0100] The data format of the combined audio data (4 channels) is:
[0101]
[0102] As described above, in this embodiment, the combined audio data combined by the first audio data and the second audio data is obtained by alternately inserting the first audio data and the second audio data.
[0103] It is worth noting that when the first audio data and the second audio data are aligned, the audio playback and the microphone collection functions are required to be started simultaneously in the read thread of the HAL layer (or the Kernel layer) so that the audio playback and the microphone collection functions start to read data, and the audio playback and the microphone collection functions are required not to occur underrun and overrun during data transmission. If overrun occurs, the buffer will be overwritten; if underrun occurs, data loss will occur; no matter whether overrun or underrun occurs, the first audio data and the second audio data will be misaligned (beat).
[0104] In addition, the combination of the aligned first audio data and the second audio data is also completed in the HAL layer (or the Kernel layer), and the AudioRecord is required to start the read thread in the AudioFlinger layer to trigger the work of the two read threads in the HAL layer, and start reading data by using the audio playback and the microphone collection functions and caching the read data in the respective Ringbuffer. It is worth noting that the audio playback and the microphone collection functions are always in a waiting state, and when the buffer corresponding to the audio playback function and the buffer corresponding to the microphone collection function both have the first amount of data, a certain amount of data is read from the start position of the buffer corresponding to the two functions respectively, and is written into the target buffer, so as to combine the aligned first audio data and the second audio data to obtain the combined audio data. For example, after 5 ms of data is in the buffer corresponding to the two functions, 5 ms of data is read from the start position of the two buffers respectively and is filled into the target buffer.
[0105] Step 14: echo cancellation based on the combined audio data.
[0106] In some embodiments, after the combined audio data is obtained, the AudioRecord (audio collection) class is used to deliver the combined audio data to the echo cancellation function, so as to perform echo cancellation on the combined audio data by using the echo cancellation function.
[0107] In other embodiments, the combined audio data can be delivered by a private interface in a callback manner.
[0108] In some embodiments, the combined audio data is subjected to echo cancellation and noise reduction processing by using an ECNR (Echo Cancellation and Noise Reduction) algorithm, so as to eliminate the corresponding first audio data in the second audio data.
[0109] Differently from the prior art, the audio data processing method provided in the application determines the time difference between starting the audio playback and starting the microphone collection, and then adjusts one of the first audio data collected by the audio playback and the second audio data collected by the microphone collection based on the time difference, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned, then the unadjusted audio data and the adjusted audio data are combined to obtain combined audio data, and echo cancellation is performed based on the combined audio data. In this way, the beat and misalignment between the first audio data and the second audio data can be solved, the two audio data can be aligned, and then the aligned two audio data can be combined to obtain the combined audio data, so that echo cancellation is performed based on the combined audio data.
[0110] Referring to Figure 5 , Figure 5 is a flowchart of a second embodiment of the audio data processing method provided in the application, and the method comprises the following steps:
[0111] Step 51: determining the time difference between starting the audio playback and starting the microphone collection by using the audio driver, and obtaining the first audio data collected by the audio playback and the second audio data collected by the microphone collection.
[0112] Step 52: determining one of the first audio data and the second audio data as the audio data to be adjusted and the other as the reference audio data based on the time difference.
[0113] Step 53: determining the number of bytes by which the audio data to be adjusted and the reference audio data differ.
[0114] Step 54: performing a data padding operation on the audio data to be adjusted based on the number of bytes, so that the adjusted audio data to be adjusted is aligned with the reference audio data.
[0115] Step 55: determining the preset format of the combined audio data.
[0116] Specifically, the preset combined audio data format is provided by a manufacturer of the audio playback and collection module and the microphone collection module, and the preset combined audio data is composed of at least 8 bytes, wherein the first byte is used to store first left channel data of the second audio data, the second byte is used to store first right channel data of the second audio data, the third byte is used to store first left channel data of the first audio data, the fourth byte is used to store first right channel data of the first audio data, the fifth byte is used to store second left channel data of the second audio data, the sixth byte is used to store second right channel data of the second audio data, the seventh byte is used to store second left channel data of the first audio data, and the eighth byte is used to store second right channel data of the first audio data.
[0117] Step 56: combining the unadjusted audio data and the adjusted audio data according to the preset combined audio data format to obtain combined audio data.
[0118] Step 57: echo cancellation based on the combined audio data.
[0119] Steps 51 to 57 can have the same or similar technical solutions as any of the above embodiments, and will not be described here.
[0120] Different from the prior art, the audio data processing method provided in the application adjusts the first audio data collected by the audio playback and collection function or the second audio data collected by the microphone collection function based on the time difference between the opening of the audio playback and collection function and the opening of the microphone collection function, so as to align the phases of the first audio data and the second audio data, and then combines the first audio data and the second audio data after phase alignment to obtain combined audio data, so that echo cancellation can be performed based on the combined audio data.
[0121] In a vehicle device application scenario, the audio playback and collection module and the microphone collection module of the vehicle device are opened, and the vehicle device also has a voice recognition function. When the voice recognition function is opened, the time difference between the opening of the audio playback and collection module and the opening of the microphone collection module needs to be determined. The first audio data collected by the audio playback and collection module or the second audio data collected by the microphone collection module is adjusted based on the time difference, so as to align the first audio data and the second audio data. The aligned first audio data and second audio data are combined to obtain combined audio data. Echo cancellation is performed based on the combined audio data.
[0122] For example, the first audio data collected by the audio playback back sampling function is "where am I", and at the same time, the user also inputs "Xiaoming Xiaoming" by voice. The microphone collection function collects the second audio data after "where am I" and "Xiaoming Xiaoming" are mixed. At this time, the time difference between the audio playback back sampling function and the microphone collection function needs to be determined. The first audio data collected by the audio playback back sampling function or the second audio data collected by the microphone collection function is adjusted based on the time difference, so that the first audio data and the second audio data are phase-aligned. The first audio data and the second audio data after phase alignment are combined to obtain combined audio data. Echo cancellation is performed based on the combined audio data. That is, "where am I" in the second audio data is eliminated to obtain "Xiaoming Xiaoming", and then it is determined whether to start voice recognition based on "Xiaoming Xiaoming".
[0123] Referring to Figure 6 , Figure 6 is a structural schematic diagram of the first embodiment of the audio data processing system provided by the present application. The audio data processing system 60 includes an audio playback back sampling module 601, a microphone collection module 602, an audio driver 603, and a processing module 604. The audio playback back sampling module 601 is configured to collect first audio data. The microphone collection module 602 is configured to collect second audio data. The audio driver 603 is connected to the audio playback back sampling module 601 and the microphone collection module 602, and is configured to determine the time difference between the audio playback back sampling function of the audio playback back sampling module 601 and the microphone collection function of the microphone collection module 602, and to obtain the first audio data collected by the audio playback back sampling module 601 and the second audio data collected by the microphone collection module 602. The processing module 604 is connected to the audio playback back sampling module 601, the microphone collection module 602, and the audio driver 603, and is configured to adjust one of the first audio data and the second audio data based on the time difference, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned, and to combine the unadjusted audio data and the adjusted audio data to obtain combined audio data, and to perform echo cancellation based on the combined audio data.
[0124] Optionally, the audio playback back sampling module 601 can be a vehicle-mounted device, a smart phone, a notebook computer, or a smart wearable device, etc. The microphone collection module 602 can be a MIC device of a sound card in a smart phone, a notebook computer, a vehicle-mounted device, etc.
[0125] In some embodiments, the audio playback sampling module 601 and the microphone sampling module 602 can be directly connected with an ADC (Analog-to-Digital Converter) or a standard I2S (Inter-IC Sound) of a main chip. At this time, the first audio data sampled by the audio playback sampling module 601 can be audio data obtained after the audio data is transmitted back to the SoC (System-on-Chip) through the I2S / ADC before being output.
[0126] In some embodiments, referring to Figure 7 , the audio data processing system 60 can further include a first buffer 701, a second buffer 702 and a third buffer 703. The first buffer 701 is connected with the audio playback sampling module 601, the audio driver 603 and the processing module 604, and is used to buffer the first audio data. The second buffer 702 is connected with the microphone sampling module 602, the audio driver 603 and the processing module 604, and is used to buffer the second audio data. The processing module 604 is used to read the first audio data from the first buffer 701, and read the second audio data from the second buffer 702. The processing module 604 is further used to adjust one of the first audio data and the second audio data based on a time difference, so that the adjusted audio data and the unadjusted audio data in the first audio data and the second audio data are phase-aligned. The processing module 604 is further used to combine the unadjusted audio data and the adjusted audio data to obtain combined audio data, and buffer the combined audio data in the third buffer 703. The third buffer 703 is connected with the processing module 604.
[0127] Referring to Figure 8 , Figure 8 is a structural schematic diagram of an embodiment of an audio data processing device provided by the present application. The audio data processing device 80 includes a memory 801 and a processor 802. The memory 801 is used to store a computer program. The processor 802 is used to execute the computer program to implement the audio data processing method of any of the above embodiments. Details are not described herein again.
[0128] Referring to Figure 9 , Figure 9 is a structural schematic diagram of an embodiment of a computer readable storage medium provided by the present application. The computer readable storage medium 90 stores a computer program 901. The computer program 901, when executed by a processor, is used to implement the audio data processing method of any of the above embodiments. Details are not described herein again.
[0129] In summary, the audio data processing method provided by the present application can solve the problem of beat and phase misalignment of two audio data, and the combined audio data obtained by combining the two aligned audio data can be used for echo cancellation.
[0130] The processor involved in the present application can be referred to as a CPU (Central Processing Unit), can be an integrated circuit chip, and can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0131] The storage medium used in the present application includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM) or an optical disk and various storage medium capable of storing program codes.
[0132] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. An audio data processing method, characterized in that, The method includes: The time difference between enabling the audio playback re-sampling function and enabling the microphone acquisition function is determined using an audio driver, and the first audio data acquired by the audio playback re-sampling function and the second audio data acquired by the microphone acquisition function are obtained; the time difference is the difference between the time when the audio playback re-sampling function is enabled and the time when the microphone acquisition function is enabled. If the time difference is negative, the second audio data is used as the audio data to be adjusted, and the first audio data is used as the reference audio data. If the difference in time is non-negative, then the first audio data is used as the audio data to be adjusted, and the second audio data is used as the reference audio data. The audio data to be adjusted is adjusted so that the adjusted audio data is phase-aligned with the reference audio data; The unadjusted audio data is combined with the adjusted audio data to obtain combined audio data; Echo cancellation is performed based on the combined audio data.
2. The method according to claim 1, characterized in that, The step of adjusting the audio data to be adjusted so that the adjusted audio data is phase-aligned with the reference audio data includes: Determine the number of bytes that differ between the audio data to be adjusted and the reference audio data; Based on the number of bytes, the audio data to be adjusted is padded to make the adjusted audio data phase-aligned with the reference audio data.
3. The method according to claim 1, characterized in that, The step of combining the unadjusted audio data with the adjusted audio data to obtain combined audio data includes: At the hardware abstraction layer, the unadjusted audio data is combined with the adjusted audio data to obtain combined audio data; or, The unadjusted audio data is combined with the adjusted audio data at the kernel level to obtain combined audio data.
4. The method according to claim 1, characterized in that, The step of combining the unadjusted audio data with the adjusted audio data to obtain combined audio data includes: Obtain a preset combined audio data format; the preset combined audio data format consists of at least the following bytes, wherein the first byte contains the first left channel data of the second audio data, the second byte contains the first right channel data of the second audio data, the third byte contains the first left channel data of the first audio data, the fourth byte contains the first right channel data of the first audio data, the fifth byte contains the second left channel data of the second audio data, the sixth byte contains the second right channel data of the second audio data, the seventh byte contains the second left channel data of the first audio data, and the eighth byte contains the second right channel data of the first audio data. The unadjusted audio data and the adjusted audio data are combined according to the preset combined audio data format to obtain combined audio data.
5. The method according to claim 4, characterized in that, Before performing echo cancellation based on the combined audio data, the method includes: passing the combined audio data to the echo cancellation function using the AudioRecord class; The echo cancellation based on the combined audio data includes: The echo cancellation function is used to cancel the echo in the combined audio data.
6. The method according to claim 1, characterized in that, Before adjusting one audio data point in the first audio data and the second audio data based on the time difference to align the phase between the adjusted and unadjusted audio data in the first and second audio data, the process includes: Obtain the adjustment parameters received by the adjustment interface, and adjust the time difference using the adjustment parameters.
7. An audio data processing system, characterized in that, The audio data processing system is used to implement the audio data processing method as described in any one of claims 1-6, and the audio data processing system includes: The audio playback and re-acquisition module is used to acquire the first audio data; Microphone acquisition module, used to acquire second audio data; An audio driver connects the audio playback back sampling module and the microphone acquisition module, used to determine the time difference between enabling the audio playback back sampling function of the audio playback back sampling module and enabling the microphone acquisition function of the microphone acquisition module, and to acquire the first audio data acquired by the audio playback back sampling module and the second audio data acquired by the microphone acquisition module. A processing module, connected to the audio playback and acquisition module, the microphone acquisition module, and the audio driver, is used to adjust one of the first and second audio data based on the time difference to achieve phase alignment between the adjusted and unadjusted audio data in the first and second audio data; combine the unadjusted audio data with the adjusted audio data to obtain combined audio data; and perform echo cancellation based on the combined audio data.
8. The audio data processing system according to claim 7, characterized in that, The audio data processing system also includes: A first buffer, connected to the audio playback and re-acquisition module, the audio driver, and the processing module, is used to buffer the first audio data; The second buffer, which is connected to the microphone acquisition module, the audio driver, and the processing module, is used to buffer the second audio data; The processing module is configured to read first audio data from the first buffer and second audio data from the second buffer, and adjust one of the first and second audio data based on the time difference to make the adjusted audio data and the unadjusted audio data in the first and second audio data phase aligned; and combine the unadjusted audio data with the adjusted audio data to obtain combined audio data, and buffer the combined audio data in a third buffer; the third buffer is connected to the processing module.
9. An audio data processing device, characterized in that, The audio data processing device includes a memory and a processor, the memory being used to store a computer program, and the processor being used to execute the computer program to implement the method as described in any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which, when executed by a processor, is used to implement the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Audio processing method and device, electronic equipment and storage medium
CN111883156A
Audio data processing method and system and storage medium
CN112509595A