Voice data acquisition method and apparatus, and device and medium
By checking and aligning the storage status of buffers in smart cars, the low-precision problem caused by unqualified voice data during speech recognition is solved, and higher speech recognition accuracy is achieved.
Patent Information
- Application Number
- PCT/CN2024/119105
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-09-14
- Publication Date
- 2025-05-08
AI Technical Summary
When the voice recognition function on smart cars is permanent, the voice recognition accuracy is low due to the unqualified voice data used during voice recognition.
By checking the storage status of the buffer, determine whether the reference and ambient sounds to be spliced are sufficient for splicing, and obtain the target voice data based on the pre-measured delay alignment and splicing.
Ensure the relative delay between the reference sound and the ambient sound in the spliced target voice data is stable, provide qualified voice data, and improve the recognition accuracy of the voice recognition function in the smart car when the background is permanent.
Smart Images

Figure CN2024119105_08052025_PF_FP_ABST
Abstract
Description
A method, device, equipment and medium for acquiring voice data
[0001] This patent application claims priority to Chinese Patent Application No. 2023114365185, filed on October 30, 2023, entitled “A Method, Apparatus, Device and Medium for Acquiring Voice Data,” and the entire contents of which are hereby incorporated by reference.
Technical field
[0002] The present application relates to the field of acoustics, more specifically to the field of computer acoustic processing, and in particular to a method, apparatus, device and medium for acquiring speech data. [Background Technology]
[0003] With the advancement of technology, cars have entered the intelligent era. Many smart cars feature intelligent voice recognition. An onboard microphone (MIC) collects ambient sound from inside the car and transmits it to the in-vehicle information system (IVI SoC). The in-vehicle information system (IVI SoC) then performs voice recognition on the received ambient sound, for example, identifying voice commands used to control the smart car.
[0004] However, the ambient sound collected by the MIC may be mixed with the sound output from the smart car's onboard speakers. For example, if the driver is giving voice commands while music is playing on the onboard speakers, the MIC will simultaneously collect the sound output from the onboard speakers and the driver's voice. The mixed sound from the speakers mixed with the ambient sound collected by the MIC can interfere with speech recognition. Therefore, the collected ambient sound needs to be processed with echo cancellation, that is, the sound output from the speakers needs to be eliminated from the collected ambient sound.
[0005] At present, the common echo cancellation solution is to collect the output audio when the car system outputs audio to the car speakers, use the collected output audio as reference sound (REF), and then use the reference sound to offset the echo in the ambient sound, thereby achieving echo cancellation processing.
[0006] [Summary of the invention]
[0007] The purpose of the embodiments of the present application is to provide a voice data acquisition method, apparatus, equipment, medium, chip and computer program product, which can, to a certain extent, solve the technical problem of low voice recognition accuracy caused by unqualified voice data used during voice recognition when the voice recognition function on a smart car is resident in the background.
[0008] A first aspect of an embodiment of the present application provides a method for acquiring voice data. The method includes: checking the storage status of a first buffer, a second buffer, a third buffer, and a fourth buffer; wherein the data in the first buffer is acquired from the third buffer, and the data in the second buffer is acquired from the fourth buffer, the third buffer is used to store reference sound collected from an audio output path, and the fourth buffer is used to store ambient sound collected from an audio input path; upon checking that the amount of data in the first buffer and the second buffer is equal and the capacities of the first buffer, the second buffer, the third buffer, and the fourth buffer are not overflowed, reading the reference sound to be spliced in the first buffer and the ambient sound to be spliced in the second buffer; aligning the read reference sound to be spliced with the ambient sound to be spliced according to a pre-determined time delay, and splicing the reference sound to be spliced with the ambient sound to be spliced after the alignment to obtain target voice data.
[0009] According to a second aspect of an embodiment of the present application, a device for acquiring speech data is provided. The device includes: a first checking module configured to check the storage status of a first buffer, a second buffer, a third buffer, and a fourth buffer; wherein the data in the first buffer is acquired from the third buffer, and the data in the second buffer is acquired from the fourth buffer; the third buffer is configured to store reference sounds collected from an audio output path, and the fourth buffer is configured to store ambient sounds collected from an audio input path; a first reading module configured to read the reference sounds to be spliced in the first buffer and the ambient sounds to be spliced in the second buffer upon checking that the amount of data in the first buffer is equal to that in the second buffer and that the capacities of the first, second, third, and fourth buffers are not overflowed; and an alignment and splicing module configured to align the read reference sounds to be spliced with the ambient sounds to be spliced based on a pre-determined time delay, and to splice the reference sounds to be spliced with the ambient sounds to be spliced after the alignment to obtain target speech data.
[0010] A third aspect of an embodiment of the present application provides an electronic device, comprising: a central processing unit, a first processor, and a memory, wherein the memory stores programs or instructions that can be run on at least one of the first processor and the central processing unit, and wherein the program or instructions contain steps of the voice data acquisition method described in the first aspect that are implemented when the program or instructions are executed by the first processor.
[0011] A fourth aspect of an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the voice data acquisition method as described in the first aspect are implemented.
[0012] A fifth aspect of an embodiment of the present application provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the voice data acquisition method as described in the first aspect.
[0013] A sixth aspect of the embodiments of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the voice data acquisition method as described in the first aspect.
[0014] In an embodiment of the present application, by checking the storage status of the first buffer, the second buffer, the third buffer, and the fourth buffer, it is determined whether the reference sound to be spliced in the first buffer is sufficient to offset the echo in the ambient sound to be spliced in the second buffer, and whether there is data loss in the two transmission processes of the reference sound collected by the reference sound acquisition hardware to the first buffer and the ambient sound collected by the ambient sound acquisition hardware to the second buffer. When it is checked that the amount of data in the first buffer is equal to that in the second buffer and the capacity of the first buffer, the second buffer, the third buffer, and the fourth buffer are not overflowed, it is determined that the reference sound to be spliced in the first buffer is sufficient to offset the echo in the ambient sound to be spliced in the second buffer, and there is no data loss in the two transmission processes. At this time, the reference sound to be spliced in the first buffer and the ambient sound to be spliced in the second buffer are read, and the read reference sound to be spliced is aligned with the ambient sound to be spliced according to a pre-determined time delay. After the alignment, the reference sound to be spliced and the ambient sound to be spliced are spliced to obtain target speech data. This ensures the stability of the relative time delay between the reference sound and the ambient sound in the target voice data obtained after splicing, that is, qualified voice data is obtained. When the target voice data is provided to the voice recognition function on the smart car for use, the reference sound and the ambient sound in the target voice data have been aligned and spliced together based on the pre-determined time delay. When performing echo cancellation processing, the reference sound can be more perfectly used to offset the echo in the ambient sound, so that the interference caused by the echo is reduced during subsequent voice recognition, thereby improving the recognition accuracy of the voice recognition function on the smart car when it is resident in the background.
Brief Description of the Drawings
[0015] FIG1 is a schematic diagram showing the working principle of an audio device in a smart car;
[0016] FIG2 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0017] FIG3 is a schematic diagram of the working principle of audio processing of the electronic device 200 shown in FIG2 ;
[0018] FIG4 is a flow chart of a method for acquiring voice data provided in an embodiment of the present application;
[0019] FIG5 is a schematic diagram of the structure of a ring buffer;
[0020] FIG6 is a flow chart of a method for acquiring voice data provided in an embodiment of the present application;
[0021] 7 is a flowchart of measuring path delay in the method for acquiring voice data provided in an embodiment of the present application;
[0022] 8 is a flowchart of some steps of measuring path delay in the voice data acquisition method provided in an embodiment of the present application;
[0023] 9 is a schematic diagram of a test result record for measuring path delay in the method for acquiring voice data provided in an embodiment of the present application;
[0024] FIG10 is a timing diagram of the data generation phase of the data acquisition method provided in the embodiment of the present application executed in the electronic device 200 provided in the embodiment of the present application;
[0025] Figures 11 and 12 are schematic diagrams of the states of the third buffer during two adjacent interruptions in the voice data acquisition method provided in an embodiment of the present application, Figure 11 is a schematic diagram of the state of the third buffer during the first interruption, and Figure 12 is a schematic diagram of the state of the third buffer during the second interruption;
[0026] 13 is a schematic diagram of the flow of voice data in the voice data acquisition method provided in an embodiment of the present application;
[0027] FIG14 is a schematic diagram of alignment and splicing of reference sound and ambient sound of voice data in the voice data acquisition method provided in an embodiment of the present application;
[0028] FIG15 is a schematic structural diagram of a device for acquiring voice data provided in an embodiment of the present application;
[0029] FIG16 is a schematic structural diagram of an electronic device also provided in an embodiment of the present application. [Specific implementation method]
[0030] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0031] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0032] Before specifically describing the voice data acquisition method, apparatus, device, storage medium, chip, and computer program product provided in the embodiments of the present application, the causes of the technical problems solved by the embodiments of the present application and the process of their discovery are described in detail.
[0033] The inventors of this application have discovered that when the voice recognition function in smart cars is running permanently in the background, the voice recognition accuracy is often low. After researching this phenomenon, the inventors found that the main cause of this low accuracy is substandard voice data used for recognition. This phenomenon and its causes are explained below using smart cars as a specific example.
[0034] FIG1 illustrates a schematic diagram of the operating principle of an audio device within a smart car. Referring to FIG1 , when the car's MIC 103 collects sound, it likely simultaneously collects music (MUSIC) output by the car's speakers 101 and speech (SOURCE) from the driver 102. The ambient sound collected by MIC 103 is a superposition of the music (MUSIC) and speech (SOURCE), including echo (MUSIC). (Hereinafter, the term "echo" refers to the sound output from the car's speakers collected by the MIC during ambient sound collection.) The collected ambient sound is then stored in a fourth buffer 105 for use by CPU 104 during speech recognition. To improve recognition accuracy, the central processing unit (CPU) 104 within the IVI SoC (not shown in FIG1 ) first reads the ambient sound MUSIC+SOURCE from the fourth buffer 105 and transfers it to the CPU 104 before recognizing the ambient sound MUSIC+SOURCE. The CPU then performs echo cancellation on the ambient sound MUSIC+SOURCE, eliminating the echo MUSIC from the ambient sound MUSIC+SOURCE and ultimately retaining the voice SOURCE. The CPU then performs voice recognition on the retained voice SOURCE and executes corresponding actions based on the recognition results, such as notifying the (intelligent vehicle) host to execute a corresponding function based on the recognized voice command.
[0035] Continuing with FIG1 , a common echo cancellation solution currently involves sampling the output audio (e.g., the music MUSIC in FIG1 ) from the audio output path. This sampled output audio is then stored in a third buffer 106 as reference audio MUSIC_ref for use by CPU 104 during echo cancellation processing. Before performing speech recognition, CPU 104 reads the reference audio MUSIC_ref from third buffer 106 and the ambient audio MUSIC+SOURCE from fourth buffer 105 and sends them to CPU 104. CPU 104 then uses the reference audio MUSIC_ref as the inverse phase of the echo MUSIC in the ambient audio MUSIC+SOURCE, thereby canceling the echo MUSIC in the ambient audio MUSIC+SOURCE, thereby achieving echo cancellation.
[0036] In the above-mentioned echo cancellation scheme, based on the knowledge of physics, it is known that for a certain output audio MUSIC, the time required for the audio MUSIC to be transmitted from the audio output path to the speaker 101, converted into sound waves by the speaker 101, transmitted to the MIC 103, and collected by the MIC 103 and then converted into audio is generally longer than the time required to directly collect the audio MUSIC from the audio output path. In other words, there is a certain time difference (referred to as the first time difference Δt1) between the first starting time point t1 of the reference sound MUSIC_ref collected from the audio output path and the second starting time point t2 of the sound corresponding to the audio MUSIC collected by the MIC 103, and the first starting time point t1 is always earlier than the second starting time point t2.
[0037] Assume that for a segment of output audio (MUSIC), the reference sound and ambient sound collection begins synchronously at time t1. From time t1 to time t3 (represented as time period T1, where t1 is before time t2 and t2 is before time t3), the reference sound (MUSIC_ref)-T1 is collected, and the ambient sound (MUSIC+SOURCE)-T1 is collected during time period T1. The audio corresponding to the time period from time t1 to time t2 in the ambient sound (MUSIC+SOURCE)-T1 does not contain an echo, while the audio corresponding to the time period from time t2 to time t3 contains an echo. Based on the foregoing, in the echo cancellation scheme described above, the reference sound MUSIC_ref serves as the inverse phase sound wave of the echo MUSIC of the ambient sound MUSIC+SOURCE, canceling the echo MUSIC. Based on this basic principle, when performing echo cancellation on the ambient sound (MUSIC+SOURCE)-T1, the reference sound (MUSIC_ref)-T1 is used as the inverse phase sound wave to cancel the echo MUSIC collected in (MUSIC+SOURCE)-T1 between time t2 and time t3. In other words, when performing echo cancellation, the first time difference Δt1 (i.e., path delay) between the reference sound and the echo collected for the same output audio needs to be considered.
[0038] Therefore, in the above-mentioned echo cancellation scheme, when the CPU 104 performs echo cancellation processing on the ambient sound (MUSIC+SOURCE)-T1 collected in a certain time period (for example, the time period T1 from time t1 to time t3), the time point when the CPU 104 obtains the reference sound (MUSIC_ref)-T1 collected in the time period T1 and the time point when the CPU 104 obtains the ambient sound (MUSIC+SOURCE)-T1 collected in the time period T1 is usually configured to have a certain second time difference Δt2 (the second time difference Δt2 is recorded as the phase difference). The relative delay Δt2 should be within a fixed range (usually the relative delay Δt2 cannot be greater than 5 milliseconds (ms), and the maximum value of the range should not exceed the aforementioned first time difference Δt1), and the relative delay Δt2 should basically remain unchanged, so that the CPU 104 can use the reference sound (MUSIC_ref)-T1 as the anti-phase sound wave of the echo MUSIC in the ambient sound (MUSIC+SOURCE)-T1 to cancel the echo MUSIC from the ambient sound (MUSIC+SOURCE)-T1, thereby achieving echo cancellation.
[0039] If the relative delay Δt2 changes significantly (for example, the relative delay is 50 ms), the CPU 104 may fail to obtain the reference sound (MUSIC_ref)-T1 in time when performing echo cancellation processing on the ambient sound (MUSIC+SOURCE)-T1. As a result, the CPU 104 lacks the inverted sound wave (MUSIC_ref)-T1 for canceling the echo MUSIC in the ambient sound (MUSIC+SOURCE)-T1 when performing echo cancellation processing on the ambient sound (MUSIC+SOURCE)-T1. As a result, the CPU 104 fails to cancel the echo MUSIC in the ambient sound (MUSIC+SOURCE)-T1 when performing echo cancellation processing on the ambient sound (MUSIC+SOURCE)-T1, that is, the echo cancellation processing fails. Therefore, when CPU104 subsequently performs voice recognition on the ambient sound (MUSIC+SOURCE)-T1 collected during the T1 time period, the ambient sound (MUSIC+SOURCE)-T1 also contains the echo MUSIC, that is, the voice data used for recognition is unqualified, resulting in the echo MUSIC also being recognized during voice recognition, interfering with the recognition of the voice data SOURCE in (MUSIC+SOURCE)-T1, resulting in low voice recognition accuracy.
[0040] The inventors of this application conducted further research on the process by which CPU 104 obtains the reference sound MUSIC_ref and the ambient sound MUSIC+SOURCE, and found that the reason why the voice data used when the voice recognition function on the smart car is resident in the background is unqualified is due to the blocking of the thread used to obtain the reference sound MUSIC_ref and the ambient sound MUSIC+SOURCE. In other words, the blocking of the thread causes the relative delay Δt2 of the reference sound MUSIC_ref and the ambient sound MUSIC+SOURCE collected in the same time period by CPU 104 to fluctuate greatly and exceed the fixed range. The following still uses the example of the relative delay of CPU 104 obtaining the reference sound (MUSIC_ref)-T1 and the ambient sound (MUSIC+SOURCE)-T1 exceeding the fixed range in FIG1 as an example to illustrate.
[0041] Please continue to refer to Figure 1. In Figure 1, there are a third buffer 106 and a fourth buffer 105 on the path of inputting the reference sound MUSIC_ref to the CPU 104 and on the path of inputting the ambient sound MUSIC+SOURCE to the CPU. Based on the principles of modern computer organization, it is known that the operating speed of the processor is much higher than that of other hardware (such as MIC103). In order to improve the operating efficiency of the processor, the processor usually does not directly read the data in the data production (or storage) hardware, but reads part of the data into the buffer for buffering. When the data currently required by the processor exists in the buffer, it is read directly from the buffer; when the data currently required by the processor does not exist in the buffer, it is read from the data production (or storage) hardware into the buffer, and then read from the buffer; thereby improving the operating efficiency of the processor (please refer to the principles of computer organization for details, which will not be repeated here). The functions of the third buffer 106 and the fourth buffer 105 in Figure 1 are similar to this. When the CPU 104 runs a speech recognition application, it needs to read the reference sound MUSIC_ref from the third buffer 106 and the ambient sound MUSIC+SOURCE from the fourth buffer 105 and send them to the CPU 104 for echo cancellation processing.
[0042] Based on the basic principles of modern computer operating systems, modern computers generally manage memory based on multiprogramming. This means that multiple programs can be loaded into memory at the same time and executed simultaneously. A running program is called a process, and a process can typically contain multiple threads. Multiple threads within a process typically execute concurrently, and for a thread to execute, it must obtain processor access. For example, in a multi-core processor, each core can schedule a thread for execution, and a processor can schedule a maximum of several threads simultaneously, depending on the number of cores. Under multiprogramming, to ensure the smooth execution of each program, threads' access to the processor must be scheduled. In computers, there are two thread scheduling models: time-sharing and preemptive. In the time-sharing model, all threads take turns obtaining processor access, and processor time slices are evenly distributed to each thread. In the preemptive model, threads with higher priority in the runnable pool are given priority for processor access. For threads with the same priority, a thread is randomly selected to occupy the processor. When it loses processor access, another thread is randomly selected to obtain processor access.
[0043] Continuing with FIG1 , when the program (or algorithm in computer terms) corresponding to the general echo cancellation scheme described above runs on CPU 104, the echo cancellation process includes at least two processing threads (referred to as a first processing thread and a second processing thread, respectively). The first processing thread activates the reference sound acquisition hardware and reads the reference sound MUSIC_ref from the third buffer 106 , while the second processing thread activates the ambient sound acquisition hardware and reads the ambient sound MUSIC+SOURCE from the fourth buffer. The first and second processing threads are scheduled according to a thread scheduling model configured for CPU 104 . When CPU 104 is overloaded, at least one of the first and second processing threads may become blocked (thread blocking can be simply understood as the thread losing processor access and temporarily or permanently ceasing execution). This can cause the relative delay Δt2 between the reference sound and the ambient sound acquired by CPU 104 during the same time period to fluctuate and exceed an allowable range (e.g., the relative delay may exceed 5ms).
[0044] For example, when CPU104 is running a speech recognition application, it notifies the first processing thread and the second processing thread to read the reference sound MUSIC_ref from the third buffer 106 and the ambient sound MUSIC+SOURCE from the fourth buffer 105 to CPU104 for echo cancellation processing, respectively. At least one of the first processing thread and the second processing thread is scheduled by the thread scheduling model to enter a blocking state such as a sleep state, a waiting state, a courtesy state or a closed state due to the high load of CPU104, resulting in the ambient sound MUSIC+SOURCE obtained by CPU104 not matching the reference sound MUSIC_ref. For example, when notifying the first processing thread to read the reference sound (MUSIC_ref)-T1 collected during time period T1 and notifying the second processing thread to read the ambient sound (MUSIC+SOURCE)-T1 collected during time period T1, the first processing thread is blocked. As a result, the ambient sound obtained by CPU 104 is (MUSIC+SOURCE)-T1, but CPU 104 fails to obtain the reference sound (MUSIC_ref)-T1 within the allowed relative delay Δt2 (e.g., within 5 ms) (CPU 104 may obtain reference sound collected in other time periods). As a result, when CPU 104 performs echo cancellation on the ambient sound (MUSIC+SOURCE)-T1, the inverted sound wave (MUSIC_ref)-T1 for canceling the echo MUSIC is missing, and the echo MUSIC in the ambient sound (MUSIC+SOURCE)-T1 cannot be eliminated. That is, the ambient sound and the reference sound used for echo cancellation do not match, and the voice data used for echo cancellation is unqualified, resulting in failure of the echo cancellation process.
[0045] After summarizing the above research process, the inventors of this application reached the following conclusions: the direct reason why the voice recognition accuracy of smart cars often decreases when the voice recognition function is running in the background is mainly: the voice data used for recognition is unqualified, resulting in the failure of echo cancellation processing, which in turn causes the voice data used for subsequent voice recognition to contain echoes, which interferes with voice recognition and leads to low voice recognition accuracy; the fundamental reason why the voice recognition accuracy often decreases when the voice recognition function is running in the background is: due to the influence of CPU load and thread scheduling model, at least one of the first processing thread that reads reference sound data into the CPU and the second processing thread that reads ambient sound data into the CPU may be blocked, resulting in a mismatch between the reference sound data read into the CPU and the ambient sound data, resulting in the CPU missing the anti-phase sound wave used to offset the echo in the ambient sound when performing echo cancellation processing on the read ambient sound data, thereby causing the echo cancellation processing to fail, which in turn causes the voice data used for subsequent voice recognition to contain echoes, which interferes with voice recognition and leads to low voice recognition accuracy.
[0046] It should be specifically stated that the explanation of the causes and discovery process of the phenomenon of low voice recognition accuracy often occurring when the voice recognition function on the smart car is resident in the background as described above should be part of the technical contribution made by the inventors of this application, and cannot be simply regarded as a cause known in the prior art or related art. Those skilled in the art usually only observe superficially that the voice recognition accuracy often occurs when the voice recognition function on the smart car is resident in the background. Without creative work (such as studying the working principle of the audio equipment on the smart car), it is impossible to find the direct cause of the low voice recognition accuracy often occurring when the voice recognition function on the smart car is resident in the background. Without further creative work (such as further studying the process of the CPU obtaining reference sound and ambient sound), it is even more impossible to find the root cause of the low voice recognition accuracy often occurring when the voice recognition function on the smart car is resident in the background (the direct cause and the root cause can be found above and will not be described here). In short, the process by which the inventors of this application discovered the direct cause and / or the root cause of the phenomenon of low voice recognition accuracy often occurring when the voice recognition function on the smart car is resident in the background is beyond the research ability and cognitive level of those skilled in the art.
[0047] In the process of inventing the technical problems described above, the inventors of this application have proposed some technical solutions that are not perfect or ideal compared to the voice data acquisition method, device, equipment, storage medium, chip and computer program product provided in the embodiments of this application. In order to facilitate understanding of the voice data acquisition method, device, equipment, storage medium, chip and computer program product provided in the embodiments of this application, some imperfect technical solutions proposed by the inventors of this application during the invention process are also explained below.
[0048] The inventors of this application have previously proposed adding a dedicated hardware circuit for echo cancellation to the device. Specifically, they designed a dedicated echo cancellation hardware circuit module. This module directly acquires a first analog signal of the reference sound collected from the audio output path and a second analog signal of the ambient sound collected from the audio input path. This module directly uses the first analog signal to suppress the second analog signal, thereby performing echo cancellation on the second analog signal. The module then converts the echo-cancelled second analog signal into a digital signal and provides it to the CPU. This avoids the problem of the CPU being blocked when acquiring the reference sound and ambient sound through two processing threads, resulting in a mismatch between the acquired reference sound and ambient sound, and thus preventing smooth echo cancellation. However, this solution requires specialized hardware engineers to design a matching echo cancellation hardware circuit. The hardware design, production, and assembly of this circuit significantly increase the hardware cost of the device.
[0049] The inventor of this application has also proposed to increase the priority of the first processing thread for obtaining the reference sound and the second processing thread for obtaining the ambient sound, so as to reduce the probability of the first processing thread and the second processing thread being blocked. Although increasing the priority of the first processing thread and the second processing thread can indeed reduce the probability of the first processing thread and the second processing thread being blocked to a certain extent, when the CPU load is too high, there is still a possibility that the first processing thread and the second processing thread will be blocked, which may cause the reference sound and the ambient sound obtained by the CPU to still not match. Therefore, the method of increasing the priority of the first processing thread and the second processing thread is not ideal for improving the recognition accuracy when the speech recognition application is resident in the background.
[0050] It should be specially stated that the above two schemes proposed by the inventor of this application are the anticipated schemes and experimental schemes in the invention process of this application. The anticipated schemes and experimental schemes have not been disclosed and should not be directly identified or regarded as the prior art of this application without evidence proving that they belong to the prior art of this application.
[0051] The voice data acquisition method, apparatus, device, storage medium, chip, and computer program product provided in the embodiments of the present application can effectively solve the above-mentioned technical problems and have better effects than the two solutions previously proposed by the inventors of this application. The following, in conjunction with the accompanying drawings, describes in detail the voice data acquisition method, apparatus, device, storage medium, chip, and computer program product provided in the embodiments of the present application.
[0052] As shown in Figure 2, a schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown. The electronic device can be a computer, a smartphone, a tablet computer, a wearable smart device, an in-vehicle information system on a chip (IVI SoC) on a smart car, or a server in a distributed system, a cloud server, an intelligent cloud computing server with artificial intelligence technology, or an intelligent cloud host. 2 , the electronic device 200 includes: a CPU 204, a first processor 207, a main memory 210, a first buffer 208, a second buffer 209, an output buffer 216 (a portion of a fixed storage area is allocated in the output buffer 216 as a third buffer 206), a fourth buffer 205, a MIC controller 211, an ADC (Analog-to-Digital Converter) device 212, a MIC 203, a speaker controller 213, an I2S (Inter-IC Sound) bus 214, a power amplifier 215, a speaker 201, and other input / output subsystems 217 (except the aforementioned ADC device 212, MIC 203, I2S bus 214, power amplifier 215, and speaker 201).
[0053] The first processor 207 is connected to the CPU 204 via a communication bus. Both the first processor 207 and the CPU 204 are connected to the main memory 210 via a bus. The first buffer 208 and the second buffer 209 are connected to the first processor 207 via a bus. The output buffer 216 (the third buffer 206) and the fourth buffer 205 are connected to the CPU 204 via a bus. In the audio input path, the MIC 203 is electrically connected to the ADC device 212. The ADC device 212 is communicatively connected to the MIC controller 211. The MIC controller 211 is communicatively connected to the CPU 204, the fourth buffer 205, and the main memory 210 via a communication bus. In the audio output path, the speaker 201 is electrically connected to the power amplifier 215. The power amplifier 215 and the speaker controller 213 are communicatively connected via an I2S bus 214. The speaker controller 213 is communicatively connected to the CPU 204, the output buffer 216 (the third buffer 206), and the main memory 210 via a communication bus.
[0054] Both the first processor 207 and the CPU 204 may include one or more processing units. Optionally, the first processor 207 and the CPU 204 may integrate an application processor and a modem processor, wherein the application processor primarily handles operations related to the operating system, user interface, and application programs, and the modem processor primarily processes communication signals, such as a baseband processor or a digital signal processor. It is understood that the modem processor may not be integrated into the first processor 207 and / or the CPU 204. In some optional embodiments, the first processor 207 may be a DSP (Digital Signal Processor), a coprocessor, or an ARM core.
[0055] The first buffer 208 , the second buffer 209 , the output buffer 216 (the third buffer 206 ), and the fourth buffer 205 are all cache memories.
[0056] The area of output buffer 216 not occupied by third buffer 206 can be used to store audio data output via the audio output path. This audio data can be transmitted to speaker controller 213 via a bus. Speaker controller 213 transmits the data to power amplifier 215 via I2S bus 214. Power amplifier 215 performs digital-to-analog conversion on the received audio data and transmits the data to speaker 201. Speaker 201 converts the received analog signal into sound for output. A first controller of output buffer 216 (not shown in FIG2 ; based on the principles of computer composition, the buffer and controller are integrated, and therefore not shown in FIG2 ; the buffer is controlled by the controller, such as controlling the reading and output of data within the buffer and communication interruption) can also capture the output audio data (e.g., creating a copy of the output audio data in third buffer 206 ) when controlling output buffer 216 to output the audio data via the audio output path to obtain a reference sound and store it in third buffer 206 . In other words, the first controller of output buffer 216 (third buffer 206 ) can function as reference sound capture hardware. After collecting the ambient sound, MIC203 obtains the corresponding analog signal and transmits it to the ADC device 212. The ADC device 212 performs analog-to-digital conversion on the received analog signal to obtain a digital signal of the collected ambient sound, and stores the digital signal of the ambient sound collected by MIC203 in the fourth buffer 205. That is, the MIC controller 211 and the second controller of the fourth buffer 205 can be used as ambient sound collection hardware (because the MIC controller 211 can store the collected data in the fourth buffer 205 almost instantly, so for the convenience of expression and explanation, the second controller of the fourth buffer 205 will be ignored when explaining the ambient sound collection hardware in the following text, and the MIC controller 211 will be directly explained as the ambient sound collection hardware).
[0057] The output buffer 216 (third buffer 206) and the fourth buffer 205 are directly controlled by the CPU 204. For example, the CPU 204 can store output audio data to be output through the audio output path in the output buffer 216; the CPU 204 can also notify the first controller and the second controller to clear the output buffer 216 (third buffer 206) and the fourth buffer 205 respectively; the CPU 204 can also monitor the interrupt status of the first controller and the second controller, and if it detects that the first controller has generated a hardware interrupt, it can notify the first processor 207 to move the data in the third buffer 206 to the first buffer 208, and if it detects that the second controller has generated a hardware interrupt, it can notify the first processor 207 to move the data in the fourth buffer 205 to the second buffer 209.
[0058] In some optional implementations, the output buffer 216 (the third buffer 206 ) and the fourth buffer 205 are DMA (Direct Memory Access) memories.
[0059] The first buffer 208 and the second buffer 209 are directly controlled by the first processor 207. For example, the first processor 207 can read data from the first buffer 208 and the second buffer 209; the first processor 207 can also notify the third controller that controls the first buffer 208 and the fourth controller that controls the second buffer 209 to clear the first buffer 208 and the second buffer 209 respectively; the first processor 207 can also transfer data in the third buffer 206 to the first buffer 208 upon receiving a first data transfer instruction from the CPU 204, and transfer data in the fourth buffer 205 to the second buffer 209 upon receiving a second data transfer instruction from the CPU 204.
[0060] In some optional embodiments, the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205 can all be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct RAM bus random access memory (DRRAM).
[0061] In some optional implementations, the first buffer 208 , the second buffer 209 , the third buffer 206 , and the fourth buffer 205 are all ring buffers. For details, please refer to the relevant description of FIG. 5 below.
[0062] The main memory 210 may be loaded with programs or instructions. The loaded programs or instructions may be programs or instructions for echo cancellation processing, speech recognition programs or instructions, programs or instructions corresponding to the voice data acquisition method provided in the embodiment of the present application, etc. The loaded programs or instructions may be executed by the processors set in the CPU 204 and the first processor 207. For example, the programs or instructions for echo cancellation processing and the speech recognition programs or instructions may be executed by the CPU 204, and the programs or instructions corresponding to the voice data acquisition method provided in the embodiment of the present application may be executed by the first processor 207. The main memory 210 may also store data. The stored data may be read by the CPU 204 and the first processor 207, and the CPU 204 and the first processor 207 may also store data in the main memory 210.
[0063] The main memory 210 can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct RAM bus random access memory (DRRAM).
[0064] Other input / output subsystems 217 include, but are not limited to, radio frequency units, network modules, audio output units, input units, sensors, display units, user input units, interface units, external memories (memory distinguished from the first buffer 208, second buffer 209, output buffer 216 / third buffer 206, fourth buffer 205 and main memory 210 in the computer system in the aforementioned electronic device) and other components. Those skilled in the art will appreciate that the electronic device 200 may also include a power supply (such as a battery) to power each component, and the power supply may be logically connected to the CPU 204 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system. The structure of the electronic device 200 shown in FIG2 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be described in detail here.
[0065] It should be understood that in an embodiment of the present application, the input unit may include a graphics processing unit (GPU), which processes image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit includes a display panel, and the display panel can be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit includes a touch panel and at least one of other input devices. A touch panel is also called a touch screen. The touch panel may include two parts: a touch detection device and a touch controller. Other input devices may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0066] The external memory can be used to store software programs and various data. The external memory may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instruction required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the external memory may include a volatile memory or a non-volatile memory, or the external memory may include both volatile and non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0067] FIG3 is a schematic diagram showing the working principle of audio processing of the electronic device 200 shown in FIG2. The components shown in FIG3 are all devices or elements in the electronic device 200 shown in FIG2, and devices or elements identified by the same reference numerals are the same (or the same type) devices or elements.
[0068] Referring to Figure 3 , when the speech recognition function on electronic device 200 is running in the background, after MIC 203 collects the sound MUSIC output by speaker 201 and the speech SOURCE spoken by person 202, the ambient sound collection hardware stores the collected ambient sound MUSIC+SOURCE in fourth buffer 205. Simultaneously, the reference sound collection hardware collects the output audio MUSIC from the audio output path to obtain a reference sound MUSIC_ref and stores it in third buffer 206.
[0069] When CPU 204 detects that the first controller has generated a hardware interrupt, CPU 204 notifies first processor 207 to move the data in third buffer 206 to first buffer 208 (i.e., CPU 204 sends a first data move instruction to first processor 207). When CPU 204 detects that the second controller has generated a hardware interrupt, CPU 204 notifies first processor 207 to move the data in fourth buffer 205 to second buffer 209 (i.e., CPU 204 sends a second data move instruction to first processor 207). When first processor 207 receives the first data move instruction, it moves the data in third buffer 206 to first buffer 208. When first processor 207 receives the second data move instruction, it moves the data in fourth buffer 205 to second buffer 209.
[0070] The first processor 207 is configured to:
[0071] Checking the storage status of the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205; wherein the data in the first buffer 208 is obtained from the third buffer 206, the data in the second buffer 209 is obtained from the fourth buffer 205, the third buffer 206 is used to store the reference sound MUSIC_ref collected from the audio output path, and the fourth buffer 205 is used to store the ambient sound MUSIC+SOURCE collected from the audio input path;
[0072] When it is checked that the amount of data in the first buffer 208 is equal to that in the second buffer 209 and the capacities of the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205 are not overflowed, reading the reference sound MUSIC_ref to be spliced in the first buffer 208 and the ambient sound MUSIC+SOURCE to be spliced in the second buffer 209;
[0073] Aligning the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced according to a pre-determined time delay, and after alignment, splicing the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced to obtain target speech data MUSIC_ref+MUSIC+SOURCE;
[0074] Send the target voice data MUSIC_ref+MUSIC+SOURCE to CPU 204;
[0075] CPU 204 is used for:
[0076] When the target voice data MUSIC_ref+MUSIC+SOURCE is received, echo cancellation processing is performed on the target voice data MUSIC_ref+MUSIC+SOURCE, that is, the reference sound MIC_ref in the target voice data MUSIC_ref+MUSIC+SOURCE is used as the inverse phase sound wave of the echo MUSIC in the target voice data MUSIC_ref+MUSIC+SOURCE to cancel the echo MUSIC in the target voice data MUSIC_ref+MUSIC+SOURCE, thereby obtaining the ambient sound SOURCE after echo cancellation processing;
[0077] The ambient sound SOURCE after the echo cancellation processing is subjected to voice recognition, and corresponding actions are executed according to the recognition result, for example, the electronic device 200 is controlled to execute corresponding functions according to the recognized voice instructions.
[0078] As shown in Figure 4, Figure 4 is a flow chart of the method for acquiring voice data provided by an embodiment of the present application. The following describes the method for acquiring voice data provided by an embodiment of the present application, taking the electronic device shown in Figures 2 and 3 as an example. However, it should be stated that this is not a specific limitation on the method for acquiring voice data provided by an embodiment of the present application. The method for acquiring voice data provided by an embodiment of the present application can also be applied to the device for acquiring voice data shown in Figure 15, the electronic device shown in Figure 16, etc. Please refer to the relevant description in the text for details. Please refer to Figures 2, 3 and 4 in combination. When the voice recognition function on the electronic device 200 is resident in the background, the first processor 207 of the electronic device 200 executes the program or instruction corresponding to the method for acquiring voice data, and implements the following steps S401 to S403:
[0079] S401 : Checking the storage status of the first buffer 208 , the second buffer 209 , the third buffer 206 , and the fourth buffer 205 .
[0080] Among them, the data in the first buffer 208 is obtained from the third buffer 206, and the data in the second buffer 209 is obtained from the fourth buffer 205. The third buffer 206 is used to store the reference sound MUSIC_ref collected by the reference sound acquisition hardware (for example, the first controller of the third buffer 206) from the audio output path, and the fourth buffer 205 is used to store the ambient sound MUSIC+SOURCE collected by the ambient sound acquisition hardware (for example, the MIC controller 211) from the audio input path.
[0081] As previously described, when the first controller of the third buffer 206 generates an interrupt, the first processor 207 moves the data in the third buffer 206 to the first buffer 208. When the second controller of the fourth buffer 205 generates an interrupt, the first processor 207 moves the data in the fourth buffer 205 to the second buffer 209. For details, please refer to the relevant description of the embodiment shown in FIG10.
[0082] In some optional implementations, the third buffer 206 and the fourth buffer 205 are DMA memories, and the interrupts generated by the first controller and the second controller, which respectively control the third buffer 206 and the fourth buffer 205, are DMA interrupts. For other related descriptions of the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205, please refer to the related descriptions of the embodiments corresponding to Figures 2 and 3, and will not be repeated here.
[0083] Checking the storage state of the buffer zone is to check the data amount of the data stored in the buffer zone and determine the capacity state of the buffer zone (ie, determine whether the capacity of the buffer zone overflows).
[0084] When the electronic device 200 turns on the voice recognition function and stays in the background, the first processor 207 checks the storage status of the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205 in real time, that is, the first processor 207 periodically or in real time checks the amount of data stored in these four buffers.
[0085] In some optional implementations, step S401 includes the following steps:
[0086] Periodically checking the amount of data in the first buffer 208 and the second buffer 209;
[0087] When it is checked that the amount of data in the first buffer 208 is equal to that in the second buffer 209 , the capacity status of the first buffer 208 , the second buffer 209 , the third buffer 206 and the fourth buffer 205 are checked.
[0088] For relevant instructions, please refer to the corresponding embodiments of the data splicing stage later in the text, which will not be repeated here.
[0089] S402: When it is checked that the amount of data in the first buffer 208 is equal to that in the second buffer 209 and the capacities of the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205 are not overflowed, the reference sound to be spliced in the first buffer 208 and the ambient sound to be spliced in the second buffer 209 are read.
[0090] Based on the foregoing, it can be seen that the first buffer 208 and the third buffer 206 both store the reference sound MUSIC_ref, while the second buffer 209 and the fourth buffer 205 both store the ambient sound MUSIC+SOURCE. The first buffer 208 and the second buffer 209 are directly accessible to the first processor 207, while the third buffer 206 and the fourth buffer 205 are not directly accessible to the first processor 207. When the first processor 207 detects that the amount of data in the first buffer 208 and the second buffer 209 is equal, it means that the reference sound MUSIC_ref to be spliced transferred from the third buffer 206 to the first buffer 208 is sufficient to offset the echo MUSIC in the ambient sound MUSIC+SOURCE to be spliced transferred from the fourth buffer 205 to the second buffer 209.
[0091] A buffer is a physical storage area used to temporarily store data while it's being moved from one location to another. When the amount of data written to a buffer exceeds its capacity, it overflows. In other words, the excess information is transferred to a buffer that doesn't have enough space. This excess information replaces the information already stored in the buffer, meaning that the data already in the buffer is overwritten by the subsequent transfer. This is illustrated below using a ring buffer as an example.
[0092] As shown in Figure 5, a schematic diagram of the structure of a ring buffer is shown. Referring to Figure 5, when writing data to the ring buffer, buffered data is written to each storage space (buffer) starting from the head of the buffer and growing clockwise toward the free space as shown in the figure. The last storage space in the buffer that stores data is generally referred to as the tail of the occupied area. For example, in the initial state, the buffer has no data yet. If N copies of data need to be written to the buffer, the first copy of the buffered data is first written to the first storage space buff1, then the second copy of the buffered data is written to the second storage space buff2, and so on until the Nth copy of the buffered data is written to the Nth storage space buffN, where N is a positive integer greater than zero. The next time data is written to the buffer, assuming the N copies of data are still in the buffer, the data is written starting from the N+1th storage space buff(N+1). When the speed of retrieving data from the buffer is slower than the speed of storing data, the tail and head of the occupied area will inevitably overlap. When the tail and head of the occupied area begin to overlap, the data already stored in the first storage space buff1 will be overwritten by subsequent information, that is, the buffer overflows, that is, the buffer is overwritten.
[0093] In some optional implementations, at least one of the first buffer 208 , the second buffer 209 , the third buffer 206 , and the fourth buffer 205 is a ring buffer as shown in FIG. 5 .
[0094] When the first processor 207 detects that the capacities of the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205 have not overflowed, this means that no overwriting has occurred in the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205. This means that no data loss has occurred during the process in which the reference sound acquisition hardware (e.g., the first controller of the third buffer 206) stores the collected reference sound in the third buffer 206, the first processor 207 transfers the reference sound from the third buffer 206 to the first buffer 208, the ambient sound acquisition hardware (e.g., the MIC controller 211) stores the collected ambient sound in the fourth buffer 205, and the first processor 207 transfers the ambient sound from the fourth buffer 205 to the second buffer 209. At this point, the first processor 207 reads the reference sound to be spliced in the first buffer 208 and the ambient sound to be spliced in the second buffer 209 into the first processor 207.
[0095] S403: Align the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced according to a pre-determined time delay, and after alignment, splice the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced to obtain target speech data MUSIC_ref+MUSIC+SOURCE.
[0096] Based on the above description of the process of discovering the relevant causes of the technical problems solved by the embodiments of the present application, it can be seen that when the voice recognition function on the electronic device 200 is resident in the background, the reference sound acquisition hardware and the ambient sound acquisition hardware are almost turned on at the same time. For an output audio MUSIC, the time point when the reference sound acquisition hardware (for example, the first controller of the third buffer 206) collects the reference sound MUSIC_ref on the output path is always earlier than the time point when the ambient sound acquisition hardware (for example, the MIC controller 211) collects the echo MUSIC, that is, the aforementioned first time difference. In other words, there is a small audio segment in the front part of the ambient sound MUSIC+SOURCE collected by the ambient sound acquisition hardware that does not have the echo MUSIC corresponding to the output audio MUSIC. Therefore, when splicing the reference sound MUSIC_ref to be spliced and the ambient sound MUSIC+SOURCE to be spliced, it is necessary to align the reference sound MUSIC_ref to be spliced with the part of MUSIC containing the echo in the ambient sound MUSIC+SOURCE to be spliced, so that the reference sound MUSIC_ref to be spliced can be spliced into the audio segment containing the echo MUSIC in the ambient sound MUSIC+SOURCE to be spliced during splicing, so as to facilitate echo cancellation processing when the target voice data MUSIC_ref+MUSIC+SOURCE is subsequently transmitted to CPU204.
[0097] Therefore, it is possible to pre-determine that, for an output audio MUSIC, the first time difference Δt1 between the time point at which the reference sound acquisition hardware (for example, the first controller of the third buffer 206) collects the reference sound MUSIC_ref on the output path is always earlier than the time point at which the ambient sound acquisition hardware (for example, the MIC controller 211) collects the echo MUSIC, that is, pre-determine the time delay Δt1, and then determine the audio segments in which the echo MUSIC does not exist and the audio segments in which the echo MUSIC exists in the ambient sound MUSIC+SOURCE to be spliced based on the pre-determined time delay, and then align the reference sound MUSIC_ref to be spliced and the audio segments in which the echo MUSIC exists in the ambient sound MUSIC+SOURCE to be spliced, and then splice the reference sound MUSIC_ref to be spliced with the audio segments in which the echo MUSIC exists in the ambient sound MUSIC+SOURCE to be spliced, to obtain the target voice data MUSIC_ref+MUSIC+SOURCE.
[0098] In some optional implementations, step S403 includes the following steps:
[0099] Splicing the reference sound data frames and the ambient sound data frames that are aligned with each other in the reference sound to be spliced and the ambient sound to be spliced into one frame to obtain a multi-frame spliced speech data frame;
[0100] The multiple frames of spliced voice data frames are spliced together according to the frame sequence of the reference sound data frames or the ambient sound data frames to obtain target voice data.
[0101] For details, please refer to the relevant description of the embodiment corresponding to Figure 14, which will not be repeated here.
[0102] In the embodiment of the present application, by checking the storage status of the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205, it is determined whether the reference sound to be spliced in the first buffer 208 is sufficient to offset the echo in the ambient sound to be spliced in the second buffer 209, and whether there is data loss in the two transmission processes of the reference sound collected by the reference sound collection hardware to the first buffer 208 and the ambient sound collected by the ambient sound collection hardware to the second buffer 209; after checking that the amount of data in the first buffer 208 is equal to that in the second buffer 209 and the first buffer 208 is equal to the amount of data in the second buffer 209, the first buffer 206 is equal to the amount of data in the second buffer 209. When the capacities of the buffer zone 208, the second buffer zone 209, the third buffer zone 206, and the fourth buffer zone 205 are not overflowed, it is determined that the reference sound to be spliced in the first buffer zone 208 is sufficient to offset the echo in the ambient sound to be spliced in the second buffer zone 209, and no data is lost in the two aforementioned transmission processes. At this time, the reference sound to be spliced in the first buffer zone 208 and the ambient sound to be spliced in the second buffer zone 209 are read, and the read reference sound to be spliced and the ambient sound to be spliced are aligned with each other according to a pre-determined time delay. After alignment, the reference sound to be spliced and the ambient sound to be spliced are spliced to obtain target speech data. This ensures the stability of the relative time delay between the reference sound and the ambient sound in the target voice data obtained after splicing, that is, qualified voice data is obtained. When the target voice data is provided to the voice recognition function on the electronic device 200 for use, since the reference sound and the ambient sound in the target voice data have been aligned and spliced together based on the pre-determined time delay, the reference sound can be more perfectly used to offset the echo in the ambient sound during echo cancellation processing, so as to reduce the interference caused by the echo during subsequent voice recognition, thereby improving the recognition accuracy of the voice recognition function on the electronic device 200 when it is resident in the background.
[0103] As shown in Figure 6, Figure 6 is a flow chart of the method for acquiring voice data provided in an embodiment of the present application. Referring to Figures 2, 3, and 6, when the voice recognition function on the electronic device 200 is resident in the background, the first processor 207 of the electronic device 200 executes the program or instructions corresponding to the method for acquiring voice data to implement the following steps S601 to S604:
[0104] S601 : Checking the storage status of the first buffer 208 , the second buffer 209 , the third buffer 206 , and the fourth buffer 205 .
[0105] S602: When it is determined that the amount of data in the first buffer 208 is equal to that in the second buffer 209 and the capacities of the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205 are not overflowing, the reference sound to be spliced in the first buffer 208 and the ambient sound to be spliced in the second buffer 209 are read.
[0106] S603: Align the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced according to a pre-determined time delay, and after alignment, splice the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced to obtain target speech data MUSIC_ref+MUSIC+SOURCE.
[0107] The above steps S601-S603 are the same as the above steps S401-S403. Please refer to the relevant instructions for details and will not be repeated here.
[0108] S604: When it is detected that the capacity of at least one of the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205 has overflowed, the first processor 207 clears the first buffer 208 and the second buffer 209, and notifies the central processing unit 204 to clear the third buffer 206 and the fourth buffer 205.
[0109] If the first processor 207 detects that any of the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205 have overflowed, this indicates that the data in the overflowed buffer may have been overwritten due to excessive load on the first processor 207 or other unexpected events, which may have prevented timely processing. This means that at least one of the collected reference sound data and the collected ambient sound may have been lost. This loss is irreversible and may not only result in unsatisfactory echo cancellation, but may also further reduce the accuracy of subsequent speech recognition. Speech recognition is performed in a timely manner, and people typically repeat their voice commands when they discover that the electronic device 200 has not correctly recognized their voice commands. Rather than using incomplete data for speech recognition, which may affect recognition accuracy, it is better to discard the incomplete data and re-collect the ambient sound before performing echo cancellation and recognition.
[0110] Therefore, the first processor 207 sends an instruction to the third controller of the first buffer 208 to clear the first buffer 208 and an instruction to the fourth controller of the second buffer 209 to clear the second buffer 209; at the same time, the first processor 207 sends a request to CPU204 to clear the third buffer 206 and the fourth buffer 205. After receiving the request, CPU204 sends an instruction to the first controller of the third buffer 206 to clear the third buffer 206 and sends an instruction to the second controller of the fourth buffer 205 to clear the fourth buffer 205; upon receiving the instructions, the third controller, the fourth controller, the first controller and the second controller respectively clear the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205.
[0111] In some embodiments, based on the embodiments corresponding to Figure 4 or Figure 6 above, when the voice recognition function on the electronic device 200 is resident in the background, the first processor 207 of the electronic device 200 executes the program or instructions corresponding to the method for obtaining the voice data, and can also implement the following steps: the first processor 207 sends the target voice data MUSIC_ref+MUSIC+SOURCE to the central processor 204.
[0112] The first processor 207 can send the target voice data MUSIC_ref+MUSIC+SOURCE to the CPU 204 via the communication bus for use by the voice recognition process running on the CPU 204. For details, please refer to the relevant description of the embodiments shown in Figures 2 and 3. This is not the focus of the embodiments of the present application and will not be repeated here.
[0113] On the basis of the above description, the following is a more detailed supplementary description of the method for acquiring voice data provided in the embodiment of the present application, still relying on the electronic device 200 shown in Figures 2 and 3.
[0114] The method for obtaining audio data provided in the embodiment of the present application generally includes the following four stages: path delay measurement stage, hardware activation stage (when the speech recognition application is running), data generation stage (when the speech recognition application is running), and data splicing stage (required by the speech recognition application).
[0115] The instructions are as follows:
[0116] Path delay measurement phase
[0117] As previously mentioned, for an output audio MUSIC, the time point at which the reference sound acquisition hardware (e.g., the first controller of the third buffer 206) captures the reference sound MUSIC_ref on the output path is always earlier than the time point at which the ambient sound acquisition hardware (e.g., the MIC controller 211) captures the echo MUSIC, i.e., the aforementioned first time difference Δt1. This first time difference Δt1 is the path delay. In other words, the path delay Δt1 describes the time difference between the time when the audio data MUSIC is captured from the audio output path and the time when the audio data MUSIC is captured from the audio input path, during the process of the target sound SOUND being captured by the audio input path and the time when the audio data MUSIC is captured from the audio output path.
[0118] The path delay Δt1 can be measured in advance for use in the data splicing stage (please refer to the relevant instructions of the data splicing stage for details). If the electronic device 200 is an IVI SoC on a smart car, its path delay can be tested and recorded in the actual vehicle calibration link. As shown in Figure 7, Figure 7 is a flow chart of measuring the path delay in the voice data acquisition method provided in an embodiment of the present application. As shown in Figure 9, Figure 9 is a schematic diagram of the test result record of measuring the path delay in the voice data acquisition method provided in an embodiment of the present application. Please refer to Figures 2, 3, 7, 8 and 9 in combination. The path delay Δt1 is measured by the first processor 207 executing the program or instruction corresponding to the voice data acquisition method to implement the following steps S701-S703:
[0119] S701: Acquire characteristic audio data.
[0120] Feature audio data is audio data with a relatively regular waveform. For example, the waveform shown in FIG9 is a triangular wave, or the waveform is a square wave. In some optional embodiments, the feature audio data includes preceding feature audio data (denoted as F-audio) and succeeding feature audio data (denoted as S-audio); wherein, during the test, the preceding feature audio data is output before the succeeding feature audio data. For example, the triangular wave shown in FIG9 is the preceding feature audio data, and the square wave is the succeeding feature audio data.
[0121] It should be noted that the characteristic audio data shown in FIG9 is merely an exemplary illustration and does not constitute a specific limitation of the present application. The characteristic audio data may also have other waveforms, such as a sine waveform. Furthermore, the output order of the preceding characteristic audio data and the subsequent characteristic audio data is not limited to that shown in FIG9 . For example, a square wave may be used as the preceding characteristic audio data, and a triangle wave may be used as the subsequent characteristic audio data.
[0122] In some optional embodiments, the characteristic audio data is pre-produced and stored in the external memory of the electronic device 200. When the first processor 207 of the electronic device 200 receives a test start command input by other input / output subsystems 217, it is loaded from the external memory to the main memory 210 and then read from the main memory 210 to the first processor 207.
[0123] In some optional implementations, the characteristic audio data is automatically generated by the first processor 207 of the electronic device 200 in response to receiving a test start command input by other input / output subsystems 217 and calling a characteristic audio data generation function.
[0124] S702: Output the characteristic audio data through the audio output path, and notify the central processor 204 to synchronously start the reference sound acquisition hardware and the ambient sound acquisition hardware to collect the audio data on the audio output path (recorded as data R) and the audio data input after the audio input path collects sound from the environment (recorded as data M).
[0125] After obtaining the characteristic audio data, the first processor 207 passes the obtained preceding characteristic audio data and subsequent characteristic audio data to the output buffer 216 in sequence, and in the process simultaneously notifies the CPU 204 to start the reference sound acquisition hardware (for example, the first controller of the third buffer 206) and the ambient sound acquisition hardware (for example, the MIC controller 211).
[0126] The first processing 207 controls the output buffer 216 to transmit the preceding characteristic audio data F-audio and the subsequent characteristic audio data S-audio to the speaker controller 213 in sequence. The speaker controller 213 transmits the preceding characteristic audio data F-audio and the subsequent characteristic audio data S-audio to the power amplifier 215 in sequence through the I2S bus 214 for digital-to-analog conversion. The power amplifier 215 transmits the analog signals corresponding to the preceding characteristic audio data F-audio and the subsequent characteristic audio data S-audio to the speaker 201 in sequence. The speaker 201 outputs the sounds F-audio-Sound and S-audio-Sound corresponding to the preceding characteristic audio data F-audio and the subsequent characteristic audio data S-audio in sequence.
[0127] The reference sound acquisition hardware (for example, the first controller of the third buffer 206 ) sequentially collects the audio data RF-audio and RS-audio corresponding to the preceding characteristic audio data F-audio and the succeeding characteristic audio data S-audio from the audio output path, and transmits them to the first processor 207 .
[0128] MIC213 collects the sounds F-audio-Sound and S-audio-Sound in sequence, obtains analog signals corresponding to the sounds F-audio-Sound and S-audio-Sound, and transmits them to the ADC device 212. The ADC device 212 performs analog-to-digital conversion on the analog signals corresponding to the sounds F-audio-Sound and S-audio-Sound in sequence, obtains digital signals MF-audio and MS-audio corresponding to the sounds F-audio-Sound and S-audio-Sound, and transmits them to the MIC controller 211. The MIC controller 211 transmits the audio data MF-audio and MS-audio collected from the audio input path to the fourth buffer 205. The second controller of the fourth buffer 205 transmits the audio data MF-audio and MS-audio to the first processor 207.
[0129] S703: Determine a path delay Δt1 according to the audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware.
[0130] After receiving the audio data RF-audio, RS-audio, MF-audio, and MS-audio, the first processor 207 determines a path delay Δt1 according to the audio data RF-audio, RS-audio, MF-audio, and MS-audio.
[0131] As shown in Figure 8, Figure 8 is a flowchart of some steps of measuring path delay in the voice data acquisition method provided in an embodiment of the present application. In some optional implementations, step S703 includes steps S801-S803:
[0132] S801: In the process of first outputting the preceding feature audio data F-audio, the proportional relationship between the preceding feature audio data RF-audio on the audio output path and the preceding feature audio data MF-audio on the audio input path is determined based on the audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware.
[0133] In the process of outputting the preceding feature audio data F-audio first, the data R and the data M are compared and identified to identify the audio data with similar waveforms in the two, that is, the collected preceding feature audio data are first identified from the data R and the data M respectively; and the first wave peak value A1 of the preceding feature audio data collected by the reference sound collection hardware and the second wave peak value A2 of the preceding feature audio data collected by the ambient sound collection hardware are respectively determined, and the preceding feature audio data RF-audio on the audio output path is determined based on the recorded first wave peak value and the second wave peak value
[0134] The audio data of the set is located in the lower half of the time axis t. In the process of outputting the pre-sequence characteristic audio data F-audio with a frequency of 1KHz (kilohertz) and a triangular wave waveform, it is recognized that the peak value of the audio data with a triangular wave waveform collected by the reference sound collection hardware is 0xeb61, and the peak value of the audio data with a triangular wave waveform collected by the ambient sound collection hardware is 0x15631, determining the pre-sequence characteristic audio on the audio output path.
[0135] The audio data collected by the hardware and the audio data collected by the ambient sound collection hardware are respectively identified one by one, and the first start and end time points (t r and t s ) and the second start and end time point (t m and t n ).
[0136] In the process of outputting the subsequent characteristic audio data S-audio, the data R and the data M are compared and identified to identify the waveforms of the two.
[0137] The hardware collects the first start and end time points (t r and t s ) and the second start and end time point (t m and t n ).
[0138] For example, please refer to FIG9 for the output frequency of 5KHz (kilohertz) and the waveform of the subsequent characteristic audio data S-audio of the square wave.
[0139] And record the first start and end time points (t r and t s ) and the second start and end time point (t m and t n ).
[0140] S803: Based on the first start and end time points (t r and t s ) and the second start and end time points (t m and t n ) to determine the path delay Δt1.
[0141] In some optional implementations, the path delay Δt1 is based on the first starting time point t r and the second starting time point t m Determine, that is, Δt1=|t r -t m |.
[0142] In some optional implementations, the path delay Δt1 is based on the first termination time point t s and the second end time point t n Determine, that is, Δt1=|t s -t n |.
[0143] Through the above steps S701 - S703 , the path delay Δt1 of the electronic device 200 can be measured.
[0144] In some optional implementations, the above steps S701-S703 may be repeated multiple times to obtain multiple (for example, M, where M is a positive integer greater than or equal to 0) path delays (Δt1)1, (Δt1)2, ..., (Δt1) M , wherein; from the multiple path delays (Δt1)1, (Δt1)2, ..., (Δt1) M A relatively stable value is determined as the final measured path delay Δt1.
[0145] Hardware startup phase
[0146] 2 and 3 , when the CPU 204 starts running the speech recognition application, the first processing thread controls the start of the reference sound acquisition hardware (e.g., the first controller of the third buffer 206 ) and the second processing thread controls the start of the ambient sound acquisition hardware (e.g., the MIC controller 211 ), and instructs the first controller to return to the first start time t aand MIC control 211 returns to the second opening time t b , and determine the first opening time t a and the second opening time t b The third time difference Δt3 = |t a -t b The third time difference Δt3 is stored in a designated storage space. For example, the designated storage space may be the main memory 210. Of course, it may be an external memory. Here, the main memory 210 is used as an example for description.
[0147] It should be noted that in modern computer system architectures, hardware operations or processing are typically prioritized throughout the system. Therefore, the timing of turning on the reference sound acquisition hardware and the ambient sound acquisition hardware is almost completely synchronized. However, to improve speech recognition accuracy, and considering the possibility (albeit a small probability) of blocking the first and second processing threads during hardware startup, the startup delay Δt3 (i.e., the aforementioned third time difference Δt3) between turning on the reference sound acquisition hardware and turning on the ambient sound acquisition hardware is also recorded. This delay Δt3 is then considered for alignment processing before subsequent speech data splicing.
[0148] Data generation stage
[0149] While CPU 204 is running the speech recognition application, it remains resident in the background. The reference sound acquisition hardware (e.g., the first controller of third buffer 206) and the ambient sound acquisition hardware (e.g., MIC controller 211) continue to operate and collect data. The detailed operation of these hardware is described in the related embodiments shown in Figures 2 and 3 and will not be further elaborated here.
[0150] It should be noted that even when no audio is output from the audio output path, the reference sound acquisition hardware (e.g., the first controller of the third buffer 206) will continue to operate and collect reference sound data. However, the collected reference sound data at this time does not contain any valid information and, in a binary bit-based data storage format, is represented by all bits being "0." For example, assuming each sample unit in the reference sound data is represented by 2 bytes and 16 bits, then even when no audio is output from the audio output path, the reference sound acquisition hardware will collect all 16 bits of the 2 bytes corresponding to each sample unit of the reference sound, i.e., each sample unit will be "0000000 0000000." Similarly, even when no audio is input from the audio input path, the ambient sound acquisition hardware (e.g., the MIC controller 211) will continue to operate and collect ambient sound data. Accordingly, the collected ambient sound data at this time does not contain any valid information and, in a binary bit-based data storage format, is represented by all bits being "0." Those skilled in the art can clearly and completely understand the above examples, and no separate examples are given here to illustrate the above.
[0151] After sampling (i.e., collecting audio), the reference sound acquisition hardware (e.g., the first controller of the third buffer 206) and the ambient sound acquisition hardware (e.g., the MIC controller 211) store the data in the third buffer 206 and the fourth buffer 205 respectively. Based on the principles of electronic information engineering, the sampling frequency of the reference sound acquisition hardware (e.g., the first controller of the third buffer 206) or the ambient sound acquisition hardware (e.g., the MIC controller 211) is usually fixed, that is, the amount of data collected within a unit time length (e.g., 1ms) is equal. In other words, within a unit time length, the amount of data written to the third buffer 206 by the reference sound acquisition hardware (e.g., the first controller of the third buffer 206) and the amount of data written to the fourth buffer 205 by the ambient sound acquisition hardware (e.g., the MIC controller 211) are equal.
[0152] At the same time, based on computer organization principles, a buffer is a high-speed buffer device with the inherent properties of a hardware interrupt. The interrupt cycle can be configured by the buffer's controller, and hardware interrupts have the highest priority in the entire computer system architecture. A computer bus is a busy "highway" used to transmit information between various hardware components. The "right of way" on this "highway" is controlled by the CPU. When an external I / O device (input / output device) needs to write data to a buffer, the I / O device sends a data write request to the buffer's controller (for a DMA buffer, this request is called a DMA request). The buffer's controller then sends a bus request to the CPU, thereby obtaining "right of way" on the "highway." When the CPU responds to the bus request and transfers the bus to the buffer's controller and the I / O device for data transfer, the I / O device begins writing data to the buffer via the bus. However, in some computer architectures, the "right of way" obtained by the buffer's controller is time-limited. Upon reaching the time limit, the controller automatically loses its "right of way" and stops writing data to the buffer, generating a hardware interrupt (for a DMA buffer, this interrupt is called a DMA interrupt) and notifying the CPU of the hardware interrupt. Therefore, for a specific buffer, the amount of data written to the buffer of the specific buffer by the controller of the specific buffer during each interrupt cycle of the specific buffer is equal. For example, the interrupt cycle of the first controller of the third buffer 206 is 3ms, that is, the first controller is interrupted once every 3ms, and the amount of data stored in the third buffer 206 during the 3ms period between each interruption is equal.
[0153] In the embodiment of the present application, the first controller of the third buffer 206 and the second controller of the fourth buffer 205 are both configured with an interruption cycle, and their working methods are the same as the working principles of the aforementioned buffers.
[0154] In some optional implementations, the first period T2 for generating an interrupt by the first controller is different from the second period T3 for generating an interrupt by the second controller.
[0155] As known from the foregoing description, when a hardware interrupt is generated by the first controller of the third buffer 206, the first processor 207 will move the data in the third buffer 206 to the first buffer 208. When a hardware interrupt is generated by the second controller of the fourth buffer 205, the first processor 207 will move the data in the fourth buffer 205 to the second buffer 209. In order to minimize the load on the first processor 207 and reduce the possibility of the processing thread being blocked, the first period T2 and the second period T3 are configured to different values. This can reduce the probability of the first processor 207 simultaneously moving data from the third buffer 206 and the fourth buffer 205, thereby reducing the load on the first processor 207.
[0156] In some optional implementations, the first period T2 and the second period T3 are not only different, but also both have a prime number of basic unit durations greater than 2.
[0157] Assuming the basic unit time length is set to T, then T2 = cT, T2 = dT, where c and d are prime numbers greater than 2, and c ≠ d. For example, if T = 1 ms, c = 3, and d = 5, then the first interrupt cycle of the first controller of the third buffer 206 is T2 = 3 ms, and the second interrupt cycle of the second controller of the fourth buffer 205 is T3 = 5 ms.
[0158] Configuring the first period T2 and the second period T3 to be prime number basic unit durations greater than 2 and different values can further reduce the probability that the first processor 207 simultaneously moves the third buffer 206 and the fourth buffer 205 .
[0159] While CPU 204 is running the speech recognition application, the speech recognition application remains resident in the background. CPU 204 and first processor 207 perform statistical encoding and transfer of the collected reference sounds based on the interrupt status of the first controller of third buffer 206 and the interrupt status of the second controller of fourth buffer 205. This is explained below.
[0160] As shown in Figure 10, Figure 10 is a timing diagram of the data generation phase of the data acquisition method provided in the embodiment of the present application executed in the electronic device 200 provided in the embodiment of the present application. In the embodiment of the present application, please refer to Figures 2, 3 and 10 in combination. The CPU 204 and the first processor 207 respectively run programs or instructions to collaboratively implement the following steps:
[0161] S1001 : CPU 204 monitors the interruption status of the first controller of the third buffer 206 and the second controller of the fourth buffer 205 .
[0162] This has been explained in the previous article and will not be repeated here.
[0163] S1002: Each time the CPU 204 detects an interruption generated by the first controller, it updates the amount of reference sound data collected according to the amount of data notified by the first controller, and sends a first data transfer instruction to the first processor 207. The first data transfer instruction carries a first data identifier and a first data amount of the data identified by the first data identifier.
[0164] The first controller notifies CPU 204 each time an interrupt is generated, and notifies CPU 204 of the end address (tail-206) of the data stored in the third buffer 206 at the time of the interrupt, as well as the amount of reference sound data (R-input) written in the previous cycle before the interrupt. CPU 204 updates the amount of reference sound data already collected, count-R-input, based on the amount of data notified by the first controller. CPU 204 encapsulates the end address (tail-206) of the data stored in the third buffer 206 at the time of the interrupt (i.e., the first data identifier) and the amount of reference sound data (R-input) written in the previous cycle before the interrupt into a first data transfer instruction and sends it to the first processor 207.
[0165] For example, the structure of the buffer shown in FIG5 is used as an example for explanation (other forms of buffers are similar). As shown in FIG11 and FIG12, FIG11 and FIG12 are schematic diagrams of the state of the third buffer during two adjacent interruptions in the voice data acquisition method provided in an embodiment of the present application. FIG11 is a schematic diagram of the state of the third buffer during the previous interruption, and FIG12 is a schematic diagram of the state of the third buffer during the next interruption (Note: Here, the first processor 207 is illustrated as an example in which no blocking occurs during the process of transferring data from the third buffer 206 to the first buffer 208). Please refer to FIG11 and FIG12, assuming that during the time period of the first interrupt cycle 0-3ms, the data written by the first controller to the third buffer is stored in the first storage space buff1 to the Nth storage space buffN shown in FIG11; during the time period of the second interrupt cycle 3ms-6ms, the data written by the first controller to the third buffer is stored in the N+1th storage space buffN+1 to the N+Xth storage space buffN+X shown in FIG12, where X is a positive integer greater than 0.
[0166] Please refer to Figure 11. When the first interrupt is generated, the first controller will notify CPU204 and notify CPU204 of the end address (tail-206) of the data stored in the third buffer 206 at the time of the interrupt - the Nth storage space buffN and the amount of reference sound data R-input written in the previous cycle before the interrupt, which is the amount of data written with a duration of 3ms (based on the aforementioned fact that the amount of data written in each interrupt cycle is equal, in this embodiment of the application, the amount of data is directly measured by duration, which will not be specifically explained later) - the amount of data stored in the first storage space buff1 to the Nth storage space buffN.
[0167] CPU 204 updates the amount of collected reference sound data, count-R-input, based on the amount of data R-input-1 notified by the first controller, and stores it in a designated storage space, such as main memory 210 shown in FIG2 . Furthermore, CPU 204 sends a first data transfer instruction to first processor 207. The first data transfer instruction carries a first data identifier (i.e., the end address (tail-206) of the data stored in third buffer 206—the Nth storage space buffN) and a first data amount of the data identified by the first data identifier (i.e., R-input = 3ms).
[0168] Please refer to Figure 12. When the second interrupt is generated for the second time, the second controller will notify CPU204 and notify CPU204 of the end address (tail-206) of the data stored in the third buffer 206 at the time of this interrupt - the N+Xth storage space buffN+X and the amount of reference sound data R-input written in the previous cycle before this interruption, which is the amount of data written with a duration of 3ms - the amount of data stored in the first storage space buffN to the N+Xth storage space buffN+X.
[0169] CPU 204 updates the collected reference sound data amount count-R-input = 3ms + 3ms = 6ms based on the data amount R-input-2 notified by the first controller, and stores it in a designated storage space, for example, in main memory 210 shown in FIG2 . Furthermore, CPU 204 sends a first data transfer instruction to first processor 207 . The first data transfer instruction carries a first data identifier (i.e., the end address (tail-206) of the data stored in third buffer 206 - N+Xth storage space buffN+X) and a first data amount of the data identified by the first data identifier (i.e., R-input = 3ms).
[0170] S1003: When the first processor 207 receives the first data transfer instruction sent by the central processing unit 204, it reads the reference sound identified by the first data identifier stored in the third buffer 206 and stores it in the first buffer 208 according to the first data identifier carried in the first data transfer instruction, and updates the amount of reference sound data that has been transferred stored in the designated storage space according to the first data amount of the data identified by the first data identifier carried in the first data transfer instruction.
[0171] After receiving the first data transfer instruction, the first processor 207 reads data between the start address (head-206) and the end address (tail-206) from the third buffer 206 based on the start address (head-206) maintained by the first processor 207 and the end address (tail-206) in the first data transfer instruction, and stores the data in the first buffer 208. Furthermore, the first processor 207 updates the amount of reference sound data (count-R) already transferred, stored in a designated storage space (e.g., the main memory 210 shown in FIG2 ), based on the first amount of data identified by the first data identifier carried in the first data transfer instruction. The start address (head-206) maintained by the first processor 207 is updated based on the end address (tail-206) of each transferred data. For example, if the data transferred this time is data stored in the first storage space buff1 to the Nth storage space buffN in the third buffer 206, after the transfer is completed, the local start address (head-206) buff0 is updated to buffN+1.
[0172] Continuing with the above example, when the first controller is interrupted for the first time, the first processor 207 moves the data stored in the first storage space buff1 to the Nth storage space buffN in the third buffer 206 to the first buffer 208, and updates the amount of reference sound data count-R that has been moved to count-R=3ms.
[0173] When the first controller is interrupted for the second time, the first processor 207 moves the data stored in the N+1th storage space buffN+1 to the N+Xth storage space buffN in the third buffer 206 to the first buffer 208, and updates the amount of reference sound data moved, count-R, to count-R=3ms+3ms=6ms.
[0174] S1004: Each time the CPU 204 detects an interruption generated by the second controller, it updates the amount of collected ambient sound data according to the amount of data notified by the second controller, and sends a second data transfer instruction to the first processor 207. The second data transfer instruction carries a second data identifier and a second data amount of the data identified by the second data identifier.
[0175] The second controller notifies CPU 204 each time an interrupt is generated, and notifies CPU 204 of the end address (tail-205) of the data stored in the fourth buffer 205 at the time of the interrupt, as well as the amount of ambient sound data written in the previous cycle before the interrupt (M-input). CPU 204 updates the amount of ambient sound data already collected, count-M-input, based on the amount of data notified by the second controller. CPU 204 encapsulates the end address (tail-205) of the data stored in the fourth buffer 205 at the time of the interrupt (i.e., the second data identifier) and the amount of ambient sound data written in the previous cycle before the interrupt (M-input) into a second data transfer instruction and sends it to the first processor 207.
[0176] The fourth buffer 205 has the same structure as the third buffer 206. The only difference between the two is the different interrupt cycles configured. This will not be illustrated with reference to the accompanying drawings, and will only be described in text below. For example, the fourth buffer 205 uses the structure of the buffer shown in FIG5 as an example for illustration (other buffers are similar). (Note: This is illustrated by taking the example of the first processor 207 not being blocked during the process of transferring data from the fourth buffer 205 to the second buffer 209). Assume that, during the first interrupt cycle of 0-5ms, the data written by the second controller to the fourth buffer 205 is stored in the first storage space buff1 to the Oth storage space buffN of the fourth buffer 205; during the second interrupt cycle of 5ms-10ms, the data written by the second controller to the fourth buffer 205 is stored in the O+1th storage space buffO+1 to the O+Yth storage space buffO+Y of the fourth buffer 205, where O and Y are positive integers greater than 0.
[0177] The second controller will notify CPU204 when an interrupt is generated for the first time, and notify CPU204 of the end address (tail-205) of the data stored in the fourth buffer 205 at the time of this interrupt - the Oth storage space buffO and the amount of ambient sound data M-input written in the previous cycle before this interruption, which is the amount of data written with a duration of 5ms - the amount of data stored in the first storage space buff1 to the Oth storage space buffO.
[0178] CPU204 updates the amount of collected ambient sound data count-M-input based on the amount of data M-input-1 notified by the second controller, and stores it in a designated storage space, for example, in the main memory 210 shown in FIG2 . Furthermore, CPU204 sends a second data transfer instruction to the first processor 207. The second data transfer instruction carries a second data identifier (i.e., the end address (tail-205) of the data stored in the fourth buffer 205 - the Oth storage space buffO) and the first data amount of the data identified by the first data identifier (i.e., R-input=3ms).
[0179] When the second interrupt occurs for the second time, the second controller will notify CPU204, and notify CPU204 of the end address (tail-205) of the data stored in the fourth buffer 205 at the time of this interruption - the O+Yth storage space buffO+Y and the amount of ambient sound data R-input written in the previous cycle before this interruption, which is the amount of data written with a duration of 5ms - the amount of data stored in the first storage space buffO to the O+Yth storage space buffO+Y.
[0180] CPU204 updates the amount of collected ambient sound data count-M-input=5ms+5ms=10ms based on the amount of data M-input-1 notified by the second controller, and stores it in a designated storage space, for example, in the main memory 210 shown in Figure 2. In addition, CPU204 sends a second data transfer instruction to the first processor 207, the second data transfer instruction carrying a second data identifier (i.e., the end address (tail-205) of the data stored in the fourth buffer 205 - the O+Y storage space buffO+Y) and the first data amount of the data identified by the first data identifier (i.e., R-input=5ms).
[0181] S1005: Upon receiving the second data transfer instruction sent by the central processing unit 204, the ambient sound identified by the second data identifier stored in the fourth buffer 205 is read and stored in the second buffer 209 according to the second data identifier carried in the second data transfer instruction, and the amount of transferred ambient sound data stored in the designated storage space is updated according to the second data amount of the data identified by the second data identifier carried in the second data transfer instruction.
[0182] After receiving the second data transfer instruction, the first processor 207 reads data between the start address (head-205) and the end address (tail-205) from the fourth buffer 205 based on the start address (head-205) maintained by the first processor 207 and the end address (tail-205) in the second data transfer instruction, and stores the data in the first buffer 208. Furthermore, the first processor 207 updates the amount of transferred ambient sound data (count-M) stored in the designated storage space (e.g., the main memory 210 shown in FIG. 2 ) based on the second amount of data identified by the second data identifier carried in the second data transfer instruction. The start address (head-205) maintained by the first processor 207 is updated based on the end address (tail-205) of each transferred data. For example, if the data transferred this time is data stored in the first storage space buff1 to the zeroth storage space buff0 in the fourth buffer 205, after the transfer is completed, the local start address (head-205) buff0 is updated to buff0+1.
[0183] Continuing with the above example, when the first controller is interrupted for the first time, the first processor 207 moves the data stored in the first storage space buff1 to the Oth storage space buffO in the fourth buffer 205 to the second buffer 209, and updates the amount of moved ambient sound data count-M to count-M=5ms.
[0184] When the first controller is interrupted for the second time, the first processor 207 moves the data stored in the O+1th storage space buffO+1 to the O+Yth storage space buffO in the fourth buffer 205 to the second buffer 209, and updates the amount of moved ambient sound data count-M to count-M=5ms+5ms=10ms.
[0185] It should be noted that each time the first controller is interrupted, the first processor 207 will move all the data in the third buffer 206 to the first buffer 208 when performing data transfer; each time the second controller is interrupted, the first processor 207 will move all the data in the fourth buffer 205 to the second buffer 209 when performing data transfer.
[0186] Data splicing stage
[0187] Based on the above, it can be seen that after the first controller and the second controller generate an interrupt, the first buffer 208 and the second buffer 209 will store the reference sound to be spliced and the ambient sound to be spliced respectively. Therefore, during the background operation of the voice recognition function of the electronic device 200, the first processor 207 also executes programs and instructions to implement the following steps S1101-S1104:
[0188] S1101 : Check the storage status of the first buffer 208 , the second buffer 209 , the third buffer 206 , and the fourth buffer 205 .
[0189] In some optional implementations, step S1101 includes the following sub-steps S11011-S11012:
[0190] S11011: Periodically check the amount of data in the first buffer 208 and the second buffer 209.
[0191] Optionally, the first processor 207 periodically reads the spliced ambient sound data amount splice-M (see the description of step S1102), the spliced reference sound data amount splice-R (see the description of step S1102), the transported ambient sound data amount count-M and the transported reference sound data amount count-R from the designated storage space (for example, the main memory 210 shown in Figure 2), checks the amount of data in the first buffer 208 and the second buffer 209, and performs an operation of subtracting the spliced reference sound data amount splice-R from the transported reference sound data amount count-R to obtain the amount of data in the first buffer 208, and performs an operation of subtracting the spliced ambient sound data amount splice-M from the transported ambient sound data amount count-M to obtain the amount of data in the second buffer 209.
[0192] S4022: When it is checked that the amount of data in the first buffer 208 is equal to that in the second buffer 209, the capacity status of the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205 are checked.
[0193] Optionally, when the first processor 207 obtains that the amount of data in the first buffer 208 is equal to the amount of data in the second buffer 209, it continues to read the amount of ambient sound collection data count-M-input and the amount of reference sound collection data count-R-input from the designated storage space (for example, the main memory 210 shown in Figure 2), and performs an operation of subtracting the amount of ambient sound collection data count-M from the amount of ambient sound data that has been transferred, to obtain the amount of data in the fourth buffer 205, and performs an operation of subtracting the amount of reference sound collection data count-R-input from the amount of reference sound data that has been transferred, to obtain the amount of data in the third buffer 206.
[0194] The first processor 207 compares the amount of data in the first buffer 208 with the capacity threshold of the first buffer 208, compares the amount of data in the second buffer 209 with the capacity threshold of the second buffer 209, compares the amount of data in the third buffer 206 with the capacity threshold of the third buffer 206, and compares the amount of data in the fourth buffer 205 with the capacity threshold of the fourth buffer 205, and determines the capacity status of the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205 based on the comparison results.
[0195] If and only if the amount of data in each of the four buffers from first buffer 208 to fourth buffer 205 is less than its own capacity threshold, it is determined that the capacity of the first buffer 208, second buffer 209, third buffer 206, and fourth buffer 205 has not overflowed. As long as the amount of data in at least one of the buffers from first buffer 208 to fourth buffer 205 is greater than or equal to its own capacity threshold, it is determined that the capacity of the buffer with the data amount greater than or equal to its own capacity threshold has overflowed.
[0196] Since first processor 207 is also a processor, it operates in the same manner as CPU 104 or CPU 204 described above, and when executing threads, it is also scheduled by the thread scheduling model configured for first processor 207. The processing thread running on first processor 207 that transfers data from third buffer 206 to first buffer 208 is the third processing thread, and the processing thread running on first processor 207 that transfers data from fourth buffer 205 to second buffer 209 is the fourth processing thread. The third and fourth processing threads' access to first processor 207 is scheduled by the thread scheduling model configured for first processor 207, meaning that the third and fourth processing threads may be blocked.
[0197] In this embodiment, the first processor 207 is pre-configured to check the period T4 of the data volume in the first buffer 208 and the second buffer 209. The first processor 207 checks the data volume in the first buffer 208 and the second buffer 209 at each T4 period point. When it is checked that the data volume in the first buffer 208 and the second buffer 209 are equal, the first processor 207 checks the data volume in the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205, thereby determining whether the capacity of at least one of the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205 overflows.
[0198] Compared with the implementation method in which the first processor 207 checks the storage status of the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205 in real time, this implementation method can more effectively reduce the load of the first processor 207, thereby reducing the probability of the third processing thread and the fourth processing thread being scheduled and blocked, and further reducing the probability of the first processor 207 being blocked in the two processes of moving data in the third buffer 206 to the first buffer 208 and moving data in the fourth buffer 205 to the second buffer 209.
[0199] S1102: When it is detected that the capacity of at least one of the first buffer 208, the second buffer 209, the third buffer 206 and the fourth buffer 205 has overflowed, the first buffer 208 and the second buffer 209 are cleared, and the central processing unit 204 is notified to clear the third buffer 206 and the fourth buffer 205.
[0200] S1103: When it is checked that the amount of data in the first buffer 208 is equal to that in the second buffer 209 and the capacities of the first buffer 208, the second buffer 209, the third buffer 206, and the fourth buffer 205 are not overflowed, the reference sound to be spliced in the first buffer 208 and the ambient sound to be spliced in the second buffer 209 are read.
[0201] Optionally, when the first processor 207 reads the reference sound to be spliced in the first buffer 208 and the ambient sound to be spliced in the second buffer 209, it updates the amount of spliced reference sound data and the amount of spliced ambient sound data stored in the designated storage space according to the data amounts corresponding to the reference sound to be spliced and the ambient sound to be spliced respectively read.
[0202] It should be noted that, during each splicing operation, the first processor 207 reads all the data in the first buffer 208 and the second buffer 209 for splicing. In the aforementioned description, during the check before each splicing operation, the first processor 207 determines the current amount of data to be spliced in the first buffer 208 and the second buffer 209. Therefore, the amount of data to be spliced determined during this check is directly used as the read data amount corresponding to the reference sound to be spliced and the ambient sound to be spliced, and is updated to the spliced reference sound data amount and the spliced ambient sound data amount stored in the designated storage space (e.g., the main memory 210 shown in FIG. 2 ).
[0203] For example, during the first splicing, the data in the current first buffer 208 and the data in the second buffer 209 each have a data volume of 30ms, then the spliced reference sound data volume splice-R stored in the main memory 210 is updated to splice-R = 30ms and the spliced ambient sound data volume splice-M is updated to splice-M = 30ms.
[0204] S1103: Align the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced according to a pre-determined time delay, and after alignment, splice the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced to obtain target speech data MUSIC_ref+MUSIC+SOURCE.
[0205] As shown in Figure 13, Figure 13 is a schematic diagram of the flow of voice data in the voice data acquisition method provided by an embodiment of the present application. As shown in Figure 14, Figure 14 is a schematic diagram of the alignment and splicing of the reference sound of the voice data and the ambient sound in the voice data acquisition method provided by an embodiment of the present application.
[0206] Please refer to FIG. 13 and FIG. 14 . In some optional implementations, the pre-measured delay includes: a pre-measured path delay Δt1; step S1103 includes:
[0207] The reference sound MUSIC_ref to be spliced and the ambient sound MUSIC+SOURCE to be spliced are aligned according to the path delay Δt1.
[0208] Referring to FIG. 13 , the collected reference sound is transferred from third buffer 206 to first buffer 208 each time a hardware interrupt is generated in third buffer 206 (e.g., the interrupt period is 3 ms). The collected ambient sound is transferred from fourth buffer 205 to second buffer 209 each time a hardware interrupt is generated in fourth buffer 205 (e.g., the interrupt period is 5 ms). Upon checking that the amount of data in first buffer 208 and second buffer 209 is equal and has not been overwritten, first processor 207 reads the reference sound (MUSIC_ref) to be spliced in first buffer 208 and the ambient sound (MUSIC+SOURCE) to be spliced in second buffer 209, aligns them, and then splices them.
[0209] Referring to Figure 14 , for example, the reference audio MUSIC_ref to be spliced contains data frames 0x10 to 0x16, and the ambient audio MUSIC+SOURCE to be spliced contains data frames 0x99 to 0x93. The reference audio and ambient audio acquisition hardware are synchronously activated at time t1, and at time t2, MIC 203 captures the echo MUSIC of the output audio MUSIC. Therefore, based on the path delay Δt1, the first data frame 0x10 of the reference audio MUSIC_ref to be spliced must be aligned with the third data frame 0x97 of the ambient audio MUSIC+SOURCE to be spliced.
[0210] In some optional implementations, the pre-determined time delay further includes: a pre-determined turn-on time delay Δt3; step S1103 includes:
[0211] According to the path delay Δt1 and the opening delay Δt3, the read reference sound to be spliced is aligned with the ambient sound to be spliced.
[0212] The reason for considering the startup delay Δt3 here has been explained above and will not be repeated here. When aligning the reference sound MUSIC_ref to be spliced with the ambient sound MUSIC+SOURCE to be spliced, it is necessary to combine the path delay Δt1 and the startup delay Δt3. For example, if the reference sound acquisition hardware is turned on later than the ambient sound acquisition hardware, the reference sound to be spliced and the ambient sound to be spliced are aligned based on the sum of the path delay Δt1 and the startup delay Δt3.
[0213] In some optional implementations, step S1103 includes the following sub-steps a and b:
[0214] Step a: splicing the reference sound data frames and the ambient sound data frames that are aligned with each other in the reference sound to be spliced and the ambient sound to be spliced into one frame to obtain a multi-frame spliced speech data frame.
[0215] As shown in Figure 14, for example, reference audio data frame 0x10 is aligned with ambient audio data frame 0x97, and the two are concatenated into one frame, resulting in concatenated voice data frame 0x97 + 0x10. Reference audio data frame 0x11 is aligned with ambient audio data frame 0x96, and the two are concatenated into one frame, resulting in concatenated voice data frame 0x96 + 0x11. Similarly, other aligned reference audio data frames and ambient audio data frames are concatenated into one frame, resulting in a multi-frame concatenated voice data frame.
[0216] Step b: splicing the multiple frames of spliced voice data frames according to the frame sequence of the reference sound data frames or the ambient sound data frames to obtain target voice data.
[0217] Please refer to Figure 14. The spliced voice data frame 0×97+0×10 is spliced with the spliced voice data frame 0×96+0×11, and the spliced voice data frame 0×96+0×11 is spliced with the next spliced voice data frame to obtain the target voice data MUSIC_ref+MUSIC+SOURCE.
[0218] For other instructions on S1101-S1104, please refer to the previous instructions and will not be repeated here.
[0219] The first processor 207 sends the target voice data MUSIC_ref+MUSIC+SOURCE to the CPU 204 .
[0220] Please refer to the previous description for details and will not be repeated here.
[0221] The voice data acquisition method provided in the embodiment of the present application can be executed by a voice data acquisition device. In the embodiment of the present application, the voice data acquisition method performed by the voice data acquisition device is taken as an example to illustrate the voice data acquisition device provided in the embodiment of the present application.
[0222] FIG15 is a block diagram of a device for acquiring voice data according to an embodiment of the present application. Referring to FIG15 , the device 120 for acquiring voice data includes:
[0223] a first checking module 121 configured to check the storage status of a first buffer, a second buffer, a third buffer, and a fourth buffer; wherein the data in the first buffer is obtained from the third buffer, the data in the second buffer is obtained from the fourth buffer, the third buffer is used to store reference sound collected from the audio output path, and the fourth buffer is used to store ambient sound collected from the audio input path;
[0224] A first reading module 122 is configured to read the reference sound to be spliced in the first buffer and the ambient sound to be spliced in the second buffer, upon detecting that the amount of data in the first buffer is equal to that in the second buffer and that the capacities of the first buffer, the second buffer, the third buffer, and the fourth buffer are not overflowed;
[0225] The alignment and splicing module 123 is configured to align the reference sound to be spliced with the ambient sound to be spliced according to a pre-determined time delay, and splice the reference sound to be spliced with the ambient sound to be spliced after alignment to obtain target speech data.
[0226] In some optional embodiments, the pre-determined delay includes: a pre-determined path delay; the path delay is used to describe the time difference between the time when the audio data is collected from the audio output path and the time when the audio data is collected from the audio input path during the process in which the audio data is output from the audio output path and becomes a target sound and the target sound is collected by the audio input path; the alignment and splicing module 123 includes:
[0227] The first alignment submodule is configured to align the reference sound to be spliced and the ambient sound to be spliced according to the path delay.
[0228] In some optional embodiments, the pre-determined delay further includes: a pre-determined startup delay; the startup delay is used to describe the time difference between the startup time of the reference sound acquisition hardware and the startup time of the ambient sound acquisition hardware; the alignment and splicing module 123 includes:
[0229] The second alignment submodule is configured to align the reference sound to be spliced with the ambient sound to be spliced according to the path delay and the opening delay.
[0230] In some optional implementations, the voice data acquisition device 120 further includes:
[0231] A first acquisition module is used to acquire characteristic audio data;
[0232] a first acquisition module, configured to output the characteristic audio data through the audio output path, and notify the central processing unit to synchronously start the reference sound acquisition hardware and the ambient sound acquisition hardware to acquire the audio data on the audio output path and the audio data inputted after the audio input path acquires the sound from the environment;
[0233] The first determining module is configured to determine a path delay based on the audio data collected by the reference sound collecting hardware and the audio data collected by the ambient sound collecting hardware.
[0234] In some optional implementations, the characteristic audio data includes preceding characteristic audio data and succeeding characteristic audio data; and the first determining module includes:
[0235] a first determining submodule for determining, in a process of outputting the preceding feature audio data, a proportional relationship between the preceding feature audio data on the audio output path and the preceding feature audio data on the audio input path based on the audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware;
[0236] A first identification submodule is configured to, in the process of subsequently outputting the subsequent feature audio data, identify the subsequent feature audio data collected by the reference sound acquisition hardware and the audio data collected by the ambient sound acquisition hardware one by one according to the ratio, and record the first start and end time points of the subsequent feature audio data collected by the reference sound acquisition hardware and the second start and end time points of the subsequent feature audio data collected by the ambient sound acquisition hardware;
[0237] The second determining submodule is configured to determine a path delay according to a time difference between the first start and end time points and the second start and end time points.
[0238] In some optional implementations, the voice data acquisition device 120 further includes:
[0239] a first execution module, configured to repeatedly execute the steps of obtaining characteristic audio data to determining a path delay based on the audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware, to obtain a plurality of path delays;
[0240] The second determining module is configured to determine a relatively stable value from the multiple path delays.
[0241] In some optional implementations, the voice data acquisition device further includes:
[0242] The first clearing module is used to clear the first buffer and the second buffer when it is detected that the capacity of at least one of the first buffer, the second buffer, the third buffer and the fourth buffer has overflowed, and to notify the central processing unit to clear the third buffer and the fourth buffer.
[0243] In some optional implementations, the first inspection module includes:
[0244] A first checking submodule, configured to periodically check the amount of data in the first buffer and the second buffer;
[0245] The second checking submodule is configured to check the capacity status of the first buffer, the second buffer, the third buffer, and the fourth buffer when it is checked that the amount of data in the first buffer is equal to that in the second buffer.
[0246] In some optional implementations, the voice data acquisition device 120 further includes:
[0247] a first transfer module configured to, upon receiving a first data transfer instruction sent by a central processing unit, read the reference sound identified by the first data identifier stored in the third buffer and store it in the first buffer according to a first data identifier carried in the first data transfer instruction, and update the amount of transferred reference sound data stored in the designated storage space according to a first data amount of the data identified by the first data identifier carried in the first data transfer instruction; wherein the first data transfer instruction is sent by the central processing unit upon detecting that an interrupt is generated by the first controller of the third buffer by the first controller;
[0248] The second transfer module is used to, upon receiving a second data transfer instruction sent by the central processing unit, read the ambient sound identified by the second data identifier stored in the fourth buffer and store it in the second buffer according to the second data identifier carried in the second data transfer instruction, and update the amount of transferred ambient sound data stored in the designated storage space according to the second data amount of the data identified by the second data identifier carried in the second data transfer instruction; wherein, the second data transfer instruction is sent by the central processing unit when it detects that the second controller in the fourth buffer of the second controller generates an interrupt.
[0249] In some optional implementations, the voice data acquisition device further includes:
[0250] A first updating module is configured to update the amount of spliced reference sound data and the amount of spliced ambient sound data stored in the designated storage space according to the read data amounts corresponding to the reference sound to be spliced and the ambient sound to be spliced;
[0251] The first inspection module includes:
[0252] a first reading submodule, configured to read the amount of collected reference sound data, the amount of collected ambient sound data, the amount of transferred reference sound data, the amount of transferred ambient sound data, the amount of spliced reference sound data, and the amount of spliced ambient sound data stored in the designated storage space; wherein the amount of collected reference sound data is updated by the central processing unit based on the amount of data counted by the first controller each time an interrupt is detected by the first controller, and the amount of collected ambient sound data is updated by the central processing unit based on the amount of data counted by the second controller each time an interrupt is detected by the second controller;
[0253] The third inspection submodule is used to check the storage status of the first buffer, the second buffer, the third buffer and the fourth buffer according to the read amount of the transported reference sound data, the amount of the spliced reference sound data, the amount of the transported ambient sound data, the amount of the spliced ambient sound data, the amount of the collected reference sound data and the amount of the collected ambient sound data.
[0254] In some optional implementations, a first period for generating an interrupt by the first controller is different from a second period for generating an interrupt by the second controller.
[0255] In some optional implementations, the first period and the second period are both prime number basic unit durations greater than 2.
[0256] In some optional implementations, the alignment and splicing module 123 includes:
[0257] A third alignment submodule is configured to splice the reference sound data frames and the ambient sound data frames that are aligned with each other in the reference sound to be spliced and the ambient sound to be spliced into one frame to obtain a multi-frame spliced speech data frame;
[0258] The first splicing submodule is configured to splice the multiple spliced voice data frames according to the frame sequence of the reference sound data frames or the ambient sound data frames to obtain target voice data.
[0259] In some optional implementations, the voice data acquisition device further includes:
[0260] The first sending module is used to send the target voice data to the central processing unit.
[0261] The voice data acquisition device 120 in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.
[0262] The voice data acquisition device 120 in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0263] The voice data acquisition device 120 provided in the embodiment of the present application can implement each process implemented in the method embodiments of Figures 2 to 14. To avoid repetition, they are not described here.
[0264] In some optional embodiments, as shown in Figure 16, an embodiment of the present application further provides an electronic device 130, including a first processor 131, a CPU 132 and a memory 133, and the memory 133 stores programs or instructions that can be run on at least one of the first processor 131 and the CPU 132. The programs or instructions stored in the memory 133 contain various steps of the above-mentioned voice data acquisition method embodiment when the program or instruction is executed by the first processor, and can achieve the same technical effect. To avoid repetition, they are not repeated here.
[0265] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0266] Any of the above-mentioned product embodiments can implement the various processes of the above-mentioned voice data acquisition method embodiment through its own processor operation, and can achieve the same technical effect. To avoid repetition, they will not be described one by one.
[0267] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned voice data acquisition method embodiment is implemented, and the same technical effect is achieved. To avoid repetition, it is not described here. The processor is the processor in the electronic device or electronic system described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0268] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned voice data acquisition method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0269] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0270] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned voice data acquisition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0271] In the embodiments provided in the examples of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0272] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of this embodiment.
[0273] In addition, each functional unit in each implementation of the embodiment of the present application may be integrated into a processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The above-mentioned integrated units may be implemented in the form of hardware or software functional units.
[0274] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0275] The above description is only an implementation method of the embodiment of the present application, and does not limit the patent scope of the embodiment of the present application. The above specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can make equivalent structural or equivalent process changes using the description and drawings of the embodiment of the present application, or directly or indirectly apply them in other related technical fields. Without departing from the scope of protection of the purpose of this application and the claims, many forms can be made, which are also included in the patent protection scope of the embodiment of the present application.
Claims
1. A method for acquiring voice data, wherein: The method comprises: Checking the storage status of the first buffer, the second buffer, the third buffer and the fourth buffer; wherein the data in the first buffer is obtained from the third buffer, the data in the second buffer is obtained from the fourth buffer, the third buffer is used to store the reference sound collected from the audio output path, and the fourth buffer is used to store the ambient sound collected from the audio input path; When it is checked that the amount of data in the first buffer is equal to that in the second buffer and the capacity of the first buffer, the second buffer, the third buffer and the fourth buffer are not overflowed, reading the reference sound to be spliced in the first buffer and the ambient sound to be spliced in the second buffer; The reference sound to be spliced is aligned with the ambient sound to be spliced according to a pre-determined time delay, and after alignment, the reference sound to be spliced is spliced with the ambient sound to be spliced to obtain target speech data.
2. The method according to claim 1, wherein: The pre-determined delay includes: a pre-determined path delay; the path delay is used to describe the time difference between the time when the audio data is collected from the audio output path and the time when the audio data is collected from the audio input path in the process in which the audio data is output from the audio output path and becomes a target sound and the target sound is collected by the audio input path; the step of aligning the reference sound to be spliced with the ambient sound to be spliced according to the pre-determined delay includes: The reference sound to be spliced and the ambient sound to be spliced that are read are aligned according to the path delay.
3. The method according to claim 2, wherein: The pre-determined delay also includes: a pre-determined start-up delay; the start-up delay is used to describe the time difference between the start-up time of the reference sound acquisition hardware and the start-up time of the ambient sound acquisition hardware; the step of aligning the reference sound to be spliced with the ambient sound to be spliced according to the pre-determined delay includes: According to the path delay and the opening delay, the read reference sound to be spliced is aligned with the ambient sound to be spliced.
4. The method according to claim 2 or 3, wherein: The path delay is measured by the following steps before checking the storage status of the first buffer, the second buffer, the third buffer, and the fourth buffer: Get feature audio data; Output the characteristic audio data through the audio output path, and notify the central processor to synchronously start the reference sound acquisition hardware and the environmental sound acquisition hardware to collect the audio data on the audio output path and the audio data input after the audio input path collects the sound from the environment; A path delay is determined based on the audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware.
5. The method according to claim 4, wherein: The characteristic audio data includes preceding characteristic audio data and succeeding characteristic audio data; The determining of the path delay according to the audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware includes: In the process of first outputting the preceding characteristic audio data, determining a proportional relationship between the preceding characteristic audio data on the audio output path and the preceding characteristic audio data on the audio input path according to the audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware; In the process of outputting the subsequent characteristic audio data, according to the ratio, the subsequent characteristic audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware are respectively identified one by one, and the first start and end time points of the subsequent characteristic audio data collected by the reference sound collection hardware and the second start and end time points of the subsequent characteristic audio data collected by the ambient sound collection hardware are recorded; The path delay is determined according to the time difference between the first start and end time points and the second start and end time points.
6. The method according to claim 4, wherein: The method further comprises: Repeating the steps of acquiring characteristic audio data to determining the path delay according to the audio data collected by the reference sound collection hardware and the audio data collected by the ambient sound collection hardware for multiple times to obtain multiple path delays; A relatively stable value is determined from the plurality of path delays.
7. The method according to claim 1, wherein: The method further comprises: When it is detected that the capacity of at least one of the first buffer, the second buffer, the third buffer and the fourth buffer has overflowed, the first buffer and the second buffer are cleared, and the central processing unit is notified to clear the third buffer and the fourth buffer.
8. The method according to claim 1 or 7, wherein: The checking of the storage status of the first buffer, the second buffer, the third buffer, and the fourth buffer includes: Periodically checking the amount of data in the first buffer and the second buffer; When it is checked that the amounts of data in the first buffer and the second buffer are equal, the capacity states of the first buffer, the second buffer, the third buffer, and the fourth buffer are checked.
9. The method according to claim 1, wherein: The method further comprises: Upon receiving a first data transfer instruction sent by the central processing unit, according to the first data identifier carried in the first data transfer instruction, the reference sound identified by the first data identifier stored in the third buffer is read and stored in the first buffer, and the amount of reference sound data that has been transferred stored in the designated storage space is updated according to the first data amount of the data identified by the first data identifier carried in the first data transfer instruction; wherein the first data transfer instruction is sent by the central processing unit when detecting that the first controller of the third buffer generates an interrupt; Upon receiving a second data transfer instruction sent by the central processing unit, the ambient sound identified by the second data identifier stored in the fourth buffer is read and stored in the second buffer according to the second data identifier carried in the second data transfer instruction, and the amount of ambient sound data that has been transferred and stored in the designated storage space is updated according to the second data amount of the data identified by the second data identifier carried in the second data transfer instruction; wherein the second data transfer instruction is sent by the central processing unit upon detecting that an interrupt is generated by the second controller of the fourth buffer.
10. The method according to claim 9, wherein: The method further comprises: According to the read data amounts corresponding to the reference sound to be spliced and the ambient sound to be spliced, respectively, updating the amount of spliced reference sound data and the amount of spliced ambient sound data stored in the designated storage space; The checking of the storage status of the first buffer, the second buffer, the third buffer, and the fourth buffer includes: Reading the amount of collected reference sound data, the amount of collected ambient sound data, the amount of moved reference sound data, the amount of moved ambient sound data, the amount of spliced reference sound data, and the amount of spliced ambient sound data stored in the designated storage space; wherein the amount of collected reference sound data is obtained by the central processing unit updating the amount of data counted by the first controller each time an interrupt is detected by the first controller, and the amount of collected ambient sound data is obtained by the central processing unit updating the amount of data counted by the second controller each time an interrupt is detected by the second controller; Check the storage status of the first buffer, the second buffer, the third buffer and the fourth buffer according to the read amount of the reference sound data that has been transported, the amount of the spliced reference sound data, the amount of the transported ambient sound data, the amount of the spliced ambient sound data, the amount of the collected reference sound data and the amount of the collected ambient sound data.
11. The method according to claim 9, wherein: A first cycle for generating interrupts by the first controller is different from a second cycle for generating interrupts by the second controller.
12. The method according to claim 11, wherein: The first period and the second period are both prime number basic unit time lengths greater than 2.
13. The method according to claim 1, wherein: After alignment, the reference sound to be spliced is spliced with the ambient sound to be spliced to obtain target speech data, including: Splicing the reference sound data frames and the ambient sound data frames in the reference sound to be spliced and the ambient sound to be spliced that are aligned with each other into one frame to obtain a multi-frame spliced speech data frame; The multiple frames of spliced voice data frames are spliced according to the frame sequence of the reference sound data frames or the ambient sound data frames to obtain target voice data.
14. The method according to claim 1, wherein: The method further comprises: The target voice data is sent to a central processing unit.
15. A voice data acquisition device, wherein: The device comprises: A first checking module is used to check the storage status of the first buffer, the second buffer, the third buffer and the fourth buffer; wherein the data in the first buffer is obtained from the third buffer, the data in the second buffer is obtained from the fourth buffer, the third buffer is used to store the reference sound collected from the audio output path, and the fourth buffer is used to store the ambient sound collected from the audio input path; A first reading module is used to read the reference sound to be spliced in the first buffer and the ambient sound to be spliced in the second buffer when checking that the amount of data in the first buffer is equal to that in the second buffer and the capacity of the first buffer, the second buffer, the third buffer and the fourth buffer are not overflowed; The alignment and splicing module is used to align the reference sound to be spliced with the ambient sound to be spliced according to the pre-determined time delay, and splice the reference sound to be spliced with the ambient sound to be spliced after alignment to obtain target voice data.
16. An electronic device, wherein: The electronic device includes: a central processing unit, a first processor and a memory, wherein the memory stores programs or instructions that can be executed on at least one of the first processor and the central processing unit, and wherein the programs or instructions contain steps of the voice data acquisition method described in any one of claims 1 to 14 that are implemented when the programs or instructions are executed by the first processor.
17. A readable storage medium, wherein: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the voice data acquisition method according to any one of claims 1 to 14 are implemented.
Citation Information
Patent Citations
Automobile with noise reduction and sound receiving device
CN112537263A
Echo cancellation method, electronic equipment and computer readable storage medium
CN113608714A
Integrated circuit and electronic device
CN113630494A
Audio data processing method and device, electronic equipment and storage medium
CN114420146A
Audio processing method and device, electronic equipment and storage medium
CN115620735A