Dynamic processing method and device for microphone delay, equipment and medium
By dynamically adjusting the microphone delay time in the vehicle cockpit system, the problem of existing technologies being difficult to meet the requirements of different audio algorithms is solved, achieving more efficient audio processing effects and stable functional performance.
Patent Information
- Application Number
- CN202510764922.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies make it difficult to effectively adjust microphone delay in vehicle-mounted cockpit systems to meet the requirements of different audio algorithms, resulting in extended development cycles and difficulty in achieving ideal audio processing effects.
By calling the recording interface in response to the target audio algorithm, the microphone delay time is dynamically adjusted to meet the requirements of the target audio algorithm. The specific method includes determining the dynamic delay time adapted to the target audio algorithm and sending a sound signal at the appropriate time to adjust the microphone delay.
The microphone delay time is dynamically adjustable to meet the requirements of different audio algorithms, reducing the manpower and time costs in the development process. It is also automatically calibrated during vehicle use to ensure functional stability.
Smart Images

Figure CN120640174A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of dynamic processing of microphone delay, and in particular to a method, apparatus, device and medium for dynamic processing of microphone delay. Background Art
[0002] To enable intelligent functions such as voice recognition and Bluetooth calls, vehicle cockpit systems are generally equipped with microphones to receive sounds. At the same time, vehicle cockpit systems are also equipped with speakers to play sounds. Therefore, in scenarios where both microphones and speakers coexist, the microphone receives the user's voice while also receiving the sound played through the speaker (this sound is often called the echo of the reference sound).
[0003] To achieve better sound reception, the sound signals received by the microphone need to be processed through audio algorithms, such as echo cancellation. These audio algorithms generally have specific requirements for the time difference between the received echo signal and the reference sound signal (usually referred to as "microphone delay"). Only when these requirements are met can better audio processing be achieved.
[0004] To meet these requirements, a common practice is to perform multiple rounds of microphone latency calibration during the development of in-vehicle cockpit systems. Based on these calibration results, the relevant hardware circuits are improved, or the speaker and microphone placement are modified. The goal is to minimize microphone latency to meet audio algorithm requirements. However, this approach is time-consuming, potentially impacting the development cycle of in-vehicle cockpit systems. Furthermore, it is difficult to meet the microphone latency requirements of various audio algorithms.
[0005] In view of this, this application is hereby filed. Summary of the Invention
[0006] The present application aims to provide a method, apparatus, device and medium for dynamic processing of microphone delay, which can make the microphone delay time dynamically adjustable to meet the different requirements of different audio algorithms for microphone delay.
[0007] In a first aspect, an embodiment of the present application provides a method for dynamically processing microphone delay, comprising:
[0008] In response to a target audio algorithm calling a recording interface, determining a dynamic delay time adapted to the target audio algorithm;
[0009] In response to the recording interface monitoring the sound signal, sending the sound signal other than the reference sound signal in the sound signal to the target audio algorithm, and waiting for the dynamic delay time before sending the reference sound signal to the target audio algorithm, so that the microphone delay time meets the requirements of the target algorithm;
[0010] In which, the reference sound signal is the original digital signal corresponding to the sound signal played by the speaker, and the microphone delay time is the time difference between the first moment and the second moment. The first moment is the moment when the target audio algorithm receives the echo signal of the reference sound signal, and the second moment is the moment when the target audio algorithm receives the reference sound signal.
[0011] According to the technical solution provided in the embodiment of the present application, optionally, determining the dynamic delay time adapted to the target audio algorithm includes:
[0012] Reading a dynamic delay time adapted to the target audio algorithm from a preset storage location;
[0013] Alternatively, a reference delay time is determined according to a requirement of the target audio algorithm on microphone delay, and a dynamic delay time adapted to the target audio algorithm is determined according to the reference delay time.
[0014] According to the technical solution provided in an embodiment of the present application, optionally, the target audio algorithm's requirement for microphone delay includes a minimum delay time and a maximum delay time, and determining the reference delay time based on the target audio algorithm's requirement for microphone delay includes:
[0015] Determining the minimum delay time as the reference delay time;
[0016] Alternatively, the maximum delay time is determined as the reference delay time;
[0017] Alternatively, the reference delay time is determined according to the minimum delay time and the maximum delay time.
[0018] According to the technical solution provided in the embodiment of the present application, optionally, determining the reference delay time according to the minimum delay time and the maximum delay time includes:
[0019] An average value of the minimum delay time and the maximum delay time is determined as the reference delay time.
[0020] According to the technical solution provided in the embodiment of the present application, optionally, determining the dynamic delay time adapted to the target audio algorithm according to the reference delay time includes:
[0021] The difference between the tested delay time and the reference delay time is determined as the dynamic delay time.
[0022] According to the technical solution provided in the embodiment of the present application, optionally, before calling the recording interface in response to the target audio algorithm and determining the dynamic delay time adapted to the target audio algorithm, the method further includes:
[0023] Controlling the loudspeaker to play a preset sound signal and monitoring the sound signal through the recording interface at the same time, wherein the sound signal monitored by the recording interface includes a reference sound signal corresponding to the preset sound signal and an echo signal of the reference sound signal;
[0024] The delay time of the test is determined according to the time when the preset audio algorithm receives the echo signal of the reference tone signal and the time when the reference tone signal is received.
[0025] According to the technical solution provided in the embodiment of the present application, optionally, determining the delay time of the test based on the moment when the echo signal of the reference tone signal is received according to a preset audio algorithm and the moment when the reference tone signal is received includes:
[0026] The difference between the time when the preset audio algorithm receives the echo signal of the reference tone signal and the time when the preset audio algorithm receives the reference tone signal is determined as the delay time of the test.
[0027] According to the technical solution provided in the embodiment of the present application, optionally, the following further comprises:
[0028] The target audio algorithm is used to eliminate the echo signal of the reference sound signal according to the reference sound signal.
[0029] In a second aspect, an embodiment of the present application further provides a device for dynamically processing microphone delay, comprising:
[0030] A determination module, configured to call a recording interface in response to a target audio algorithm and determine a dynamic delay time adapted to the target audio algorithm;
[0031] a sending module, configured to send, in response to the recording interface monitoring a sound signal, a sound signal other than a reference sound signal in the sound signal to the target audio algorithm, and to wait for the dynamic delay time before sending the reference sound signal to the target audio algorithm, so that the microphone delay time meets the requirements of the target algorithm;
[0032] In which, the reference sound signal is the original digital signal corresponding to the sound signal played by the speaker, and the microphone delay time is the time difference between the first moment and the second moment. The first moment is the moment when the target audio algorithm receives the echo signal of the reference sound signal, and the second moment is the moment when the target audio algorithm receives the reference sound signal.
[0033] In a third aspect, an embodiment of the present application further provides an electronic device, comprising:
[0034] processor and memory;
[0035] The processor is configured to execute the steps of the method for dynamically processing microphone delay as described in any one of the embodiments by calling the program or instructions stored in the memory.
[0036] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute the steps of the dynamic processing method of microphone delay as described in any embodiment.
[0037] In summary, the present application proposes a dynamic processing method for microphone delay, specifically, determining an adaptive dynamic delay time for different audio algorithms. For example, when the target audio algorithm calls the recording interface, the dynamic delay time adapted to the target audio algorithm is determined, and then the sound signals except the reference sound signal in the sound signals monitored by the recording interface are sent to the target audio algorithm, and the reference sound signal is sent to the target audio algorithm after waiting for the dynamic delay time, thereby achieving the purpose of making the microphone delay time meet the requirements of the target algorithm, wherein the microphone delay time is the time difference between the first moment and the second moment, the first moment is the moment when the target audio algorithm receives the echo signal of the reference sound signal, and the second moment is the moment when the target audio algorithm receives the reference sound signal. The purpose of making the microphone delay time dynamically adjustable is achieved, and the different requirements of different audio algorithms for microphone delay can be met. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a process of a dynamic processing method for microphone delay provided by an embodiment of the present application Figure 1 ;
[0039] Figure 2 is a schematic diagram of microphone delay provided by an embodiment of the present application;
[0040] Figure 3 This is a process of a dynamic processing method for microphone delay provided by an embodiment of the present application Figure 2 ;
[0041] Figure 4 1 is a schematic structural diagram of a dynamic processing device for microphone delay provided in an embodiment of the present application;
[0042] Figure 5 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.
[0044] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0045] Figure 1 This is a flow chart of a method for dynamically processing microphone delay provided by an embodiment of the present application. Figure 1 The dynamic processing method of microphone delay specifically includes the following steps:
[0046] S110 : In response to a target audio algorithm calling a recording interface, determining a dynamic delay time adapted to the target audio algorithm.
[0047] Typically, in real-time voice interaction scenarios, users can trigger the target audio algorithm to call the recording interface by saying a wake-up word, so as to monitor sound signals through the recording interface. The sound signals include the voice signal emitted by the user and the ambient sound signal (for example, car horns, vendors' hawking, etc. are ambient sounds). Furthermore, the ambient sound signal may also include the sound signal played by the speaker. For example, the in-car navigation application will randomly play navigation prompts (for example, please go straight at the intersection ahead, do not go uphill, or there is a red light running capture ahead, please drive according to traffic rules, etc.).
[0048] The sound signal played by the speaker is usually defined as the echo signal of the reference sound, and the original digital signal corresponding to the sound signal played by the speaker is usually defined as the reference sound signal. In other words, the reference sound signal is the sound digital signal output by the digital signal processor before it is sent to the sound playback channel.
[0049] In application scenarios such as video conferencing, Internet phone or intercom, when the user turns on the "speak" or "talk" function, the target audio algorithm is triggered to call the recording interface.
[0050] In the offline recording scenario, when the user presses the "Start Recording" button, the target audio algorithm is triggered to call the recording interface.
[0051] The target audio algorithms built into different applications are not the same. The target audio algorithm can specifically be an echo cancellation algorithm, and different echo cancellation algorithms have different requirements for microphone delay. The target audio algorithm's requirements for microphone delay include a minimum delay time and a maximum delay time, where the minimum delay time and the maximum delay time both represent a specific duration, for example, the minimum delay time is 5ms and the maximum delay time is 8ms. The delay time means: the difference between the time when the target audio algorithm receives the reference tone signal and the time when the echo signal of the reference tone is received, or in other words, the time when the echo signal arrives at the target audio algorithm later than the reference tone signal. This time cannot exceed 8ms, otherwise a good echo cancellation effect will not be achieved.
[0052] The typical echo cancellation principle involves constructing an echo path model based on the delay time. The echo signal is assumed to be the reference sound signal s(t) delayed by τ and attenuated by α, i.e., the echo signal e(t) = α·s(t-τ). This generates a reverse cancellation signal, multiplying the estimated echo signal e(t) by -1 to obtain a signal -e(t) with the opposite phase to the echo signal. The microphone signal (including the user's voice and the echo) is then added to the reverse cancellation signal to eliminate the echo from the microphone signal.
[0053] Therefore, in order to achieve better echo cancellation effect, the reference audio signal and the echo signal of the reference audio signal should be transmitted to the target audio algorithm in a timely manner according to the requirements of the target audio algorithm for microphone delay. It should be ensured that the echo signal does not arrive at the target audio algorithm too late, or that the reference audio signal does not arrive at the target audio algorithm too early, that is, it should be ensured that the difference between the time when the echo signal arrives at the target audio algorithm and the time when the reference audio signal arrives at the target audio algorithm does not exceed the maximum delay time required by the target audio algorithm.
[0054] In some embodiments, determining the dynamic delay time adapted to the target audio algorithm includes: reading the dynamic delay time adapted to the target audio algorithm from a preset storage location; that is, the dynamic delay time adapted to the target audio algorithm is predetermined and can be directly read from the designated storage location, which is conducive to improving overall efficiency and ensuring real-time performance.
[0055] In other embodiments, determining a dynamic delay time adapted to the target audio algorithm includes: determining a reference delay time based on the target audio algorithm's requirements for microphone delay, and determining a dynamic delay time adapted to the target audio algorithm based on the reference delay time.
[0056] Furthermore, after determining the dynamic delay time adapted to the target audio algorithm based on the reference delay time, the dynamic delay time is associated with the target audio algorithm and stored, so that when the target audio algorithm calls the recording interface again, the dynamic delay time adapted to the target audio algorithm can be directly read without determining the dynamic delay time online each time, so as to improve overall efficiency and ensure real-time performance.
[0057] Among them, the reference delay time is determined according to the requirement of the target audio algorithm for microphone delay, including: determining the minimum delay time required by the target audio algorithm as the reference delay time; or, determining the maximum delay time required by the target audio algorithm as the reference delay time; or, determining the reference delay time according to the minimum delay time and the maximum delay time, for example, determining the average of the minimum delay time and the maximum delay time as the reference delay time.
[0058] The reference delay time means the difference between the time when the echo signal reaches the target audio algorithm and the time when the reference audio signal reaches the target audio algorithm.
[0059] Understandably, Figure 2 As shown, the reference sound signal is the original digital signal corresponding to the sound signal played by the loudspeaker, or in other words, the reference sound signal is the sound digital signal output by the digital signal processor and has not yet been sent to the sound playback channel. The reference sound signal is sent directly to the audio algorithm through the link, and the time it takes to transmit the reference sound signal to the audio algorithm is recorded as T2. The echo signal is picked up by the microphone. Specifically, the loudspeaker first plays the sound signal, and then the played sound signal is picked up by the microphone to obtain the echo signal. This period includes the time it takes for the sound signal to propagate in the air, T0, and then the echo signal is transmitted to the audio algorithm through the link between the microphone node and the audio algorithm. This period may also involve the time it takes to convert the analog signal picked up by the microphone into a digital signal, and some may also include the time it takes for basic processing operations such as signal denoising or enhancement, as well as the time it takes for data transmission (these total times correspond to Figure 2 In short, the reference sound signal will reach the audio algorithm first, while the echo signal will arrive at the audio algorithm later. The time difference between the two can be expressed as △T=T0+T1-T2. In systems developed by different manufacturers and in different application scenarios, the time required to transmit the echo signal to the audio algorithm through the link between the microphone node and the audio algorithm (i.e., T1) is not fixed, and it is difficult to change the link time. In addition, different audio algorithms have different requirements for microphone delay. Therefore, the embodiment of the present application proposes a method for changing the microphone delay through software methods, which can meet the different requirements of different audio algorithms for microphone delay and provide a guarantee for obtaining a better echo cancellation effect.
[0060] Furthermore, determining a dynamic delay time adapted to the target audio algorithm according to the reference delay time includes: determining a difference between the tested delay time and the reference delay time as the dynamic delay time.
[0061] The test delay time refers to the time difference between transmitting the echo signal and the reference tone signal to the audio algorithm along the existing transmission link (including the transmission link between the reference tone signal and the audio algorithm and the transmission link between the echo signal (i.e., the microphone signal) and the audio algorithm). In other words, the test delay time is the difference between the moment the audio algorithm receives the echo signal of the reference tone signal and the moment the preset audio algorithm receives the reference tone signal.
[0062] The difference between the tested delay time and the reference delay time is then determined as the dynamic delay time. Specifically, because the tested delay time is greater than the reference delay time, that is, the time difference between the time the reference tone signal arrives at the audio algorithm and the time the echo signal arrives at the audio algorithm exceeds the minimum delay time required by the audio algorithm, the requirements of the audio algorithm are not met. Therefore, it is necessary to change the time the reference tone signal arrives at the audio algorithm and / or the time the echo signal arrives at the audio algorithm. To make the time the echo signal arrives at the audio algorithm faster or shorter, it may be necessary to replace the hardware processor or transmission link, which is complex, difficult to implement, and costly. Therefore, this embodiment changes the time difference between the reference tone signal and the echo signal by delaying the time the reference tone signal arrives at the audio algorithm, thereby meeting the requirements of the audio algorithm. Specifically, in response to the recording interface monitoring the sound signal, the sound signal other than the reference tone signal is sent to the target audio algorithm, and the reference tone signal is sent to the target audio algorithm after the dynamic delay time has elapsed. That is, the software method is used to reduce the difference between the time when the echo signal reaches the target audio algorithm and the time when the reference sound signal reaches the target audio algorithm, thereby meeting the microphone delay requirement of the target audio algorithm.
[0063] S120. In response to the recording interface monitoring the sound signal, the sound signal other than the reference sound signal is sent to the target audio algorithm, and the reference sound signal is sent to the target audio algorithm after waiting for the dynamic delay time, so that the microphone delay time meets the requirements of the target algorithm.
[0064] The embodiment of the present application provides a dynamic processing method for microphone delay, which determines a reference delay time according to the requirements of the target audio algorithm for microphone delay; determines a dynamic delay time adapted to the target audio algorithm according to the reference delay time; then sends the sound signals except the reference sound signal in the sound signals monitored by the recording interface to the target audio algorithm, and waits for the dynamic delay time before sending the reference sound signal to the target audio algorithm, thereby achieving the purpose of making the microphone delay time meet the requirements of the target algorithm, wherein the microphone delay time is the time difference between the first moment and the second moment, the first moment being the moment when the target audio algorithm receives the echo signal of the reference sound signal, and the second moment being the moment when the target audio algorithm receives the reference sound signal. The purpose of making the microphone delay time dynamically adjustable is achieved, and the different requirements of different audio algorithms for microphone delay can be met.
[0065] In some implementations, before the method of calling the recording interface in response to the target audio algorithm and determining the dynamic delay time adapted to the target audio algorithm further includes:
[0066] Control the speaker to play a preset sound signal while monitoring the sound signal through the recording interface. The sound signal monitored by the recording interface includes a reference sound signal corresponding to the preset sound signal and an echo signal of the reference sound signal. Specifically, for example, if the speaker plays the preset sound signal "Please note that you will be photographed if you run a red light," the reference sound signal corresponding to the preset sound signal is the original digital signal output by the digital processor corresponding to "Please note that you will be photographed if you run a red light," while the echo signal of the reference sound signal refers to the sound signal "Please note that you will be photographed if you run a red light" picked up by the microphone.
[0067] The test delay time is determined based on the time when the preset audio algorithm receives the echo signal of the reference tone signal and the time when the preset audio algorithm receives the reference tone signal. Specifically, the test delay time is determined as the difference between the time when the preset audio algorithm receives the echo signal of the reference tone signal and the time when the preset audio algorithm receives the reference tone signal. The preset audio algorithm can be a target audio algorithm or another audio algorithm, and is not limited in this embodiment.
[0068] Based on the above embodiments, Figure 3 The flowchart of dynamic processing of microphone delay shown includes the following steps:
[0069] S1. Control the speaker to play the calibration audio.
[0070] To calculate the test latency, or to test inherent microphone latency, first play a sound through the speaker. This sound can be a dedicated calibration sound or sounds typically produced during vehicle system use, such as startup music, keypad sounds, or alarms. This sound should be characteristic enough to be easily distinguished from the subsequently recorded microphone signal and reference sound.
[0071] S2, microphone recording and monitoring reference sound signal.
[0072] While the speaker plays the calibration audio, the microphone picks up the sound signal, which includes the echo signal of the reference sound signal for subsequent comparison.
[0073] S3. Calculate the test delay time.
[0074] The test delay time is calculated based on the time when the reference tone signal reaches the specified audio algorithm and the time when the echo signal of the reference tone signal reaches the specified audio algorithm.
[0075] S4. The target audio algorithm calls the recording interface.
[0076] S5. Determine a dynamic delay time adapted to the target audio algorithm according to requirements of the target audio algorithm and the tested delay time.
[0077] S6. Associate the dynamic delay time with the target audio algorithm and store it.
[0078] S7. In response to the recording interface monitoring the sound signal, sending the sound signal other than the reference sound signal to the target audio algorithm, and waiting for the dynamic delay time before sending the reference sound signal to the target audio algorithm.
[0079] S8. The target audio algorithm calls the recording interface.
[0080] S9. Read the dynamic delay time adapted to the target audio algorithm.
[0081] Subsequently, when other audio algorithms call the recording interface, S5-S6 are re-executed, and the adapted dynamic delay time is recalculated according to the requirements of the current audio algorithm, so that the microphone delay can be dynamically calibrated according to different audio algorithms.
[0082] This method significantly reduces the time and effort required to calibrate microphone delay during development, accelerating project progress. Furthermore, after a vehicle is released to the market, automatic calibration can be implemented to prevent malfunctions if microphone delay changes due to various factors.
[0083] This embodiment proposes a dynamic delay of the reference sound, which is calculated based on the current microphone delay and the requirements of the target audio algorithm and can be dynamically adjusted according to the requirements of different audio algorithms.
[0084] This solution calibrates microphone latency by dynamically delaying a reference audio signal, ensuring it meets the requirements of different audio algorithms. This reduces the time and effort required to perform multiple rounds of microphone latency calibration during development, shortening project schedules. After vehicle launch, microphone latency remains stable despite changes in installation location, system version, and other factors, ensuring proper functionality. This solution also meets the varying microphone latency requirements of different audio algorithms.
[0085] Based on the same inventive concept, corresponding to any of the above embodiments and methods, the present application also provides a dynamic processing device for microphone delay, such as Figure 4 Shown, including:
[0086] A determination module 410 is used to call a recording interface in response to a target audio algorithm and determine a dynamic delay time adapted to the target audio algorithm; a sending module 420 is used to send sound signals other than a reference sound signal in the sound signal to the target audio algorithm in response to the recording interface monitoring a sound signal, and to wait for the dynamic delay time before sending the reference sound signal to the target audio algorithm so that the microphone delay time meets the requirements of the target algorithm; wherein the reference sound signal is the original digital signal corresponding to the sound signal played by the speaker, and the microphone delay time is the time difference between a first moment and a second moment, the first moment being the moment when the target audio algorithm receives the echo signal of the reference sound signal, and the second moment being the moment when the target audio algorithm receives the reference sound signal.
[0087] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0088] The apparatus of the above embodiment is used to implement the corresponding method for dynamically processing microphone delay in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0089] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0090] The apparatus of the above embodiment is used to implement the corresponding method for dynamically processing microphone delay in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0091] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown, the electronic device 500 includes one or more processors 501 and a memory 502 .
[0092] The processor 501 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 500 to perform desired functions.
[0093] The memory 502 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 501 may execute the program instructions to implement the dynamic processing method for microphone delay of any embodiment of the present application described above and / or other desired functions. Various contents such as initial external parameters, thresholds, etc. may also be stored in the computer-readable storage medium.
[0094] In one example, the electronic device 500 may further include an input device 503 and an output device 504, which are interconnected via a bus system and / or other connection mechanisms (not shown). The input device 503 may include, for example, a keyboard, a mouse, etc. The output device 504 may output various information to the outside, including warning information, braking force, etc. The output device 504 may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto.
[0095] Of course, to simplify, Figure 5 Only some of the components related to the present application in the electronic device 500 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 500 may further include any other appropriate components according to specific application scenarios.
[0096] In addition to the above methods and devices, embodiments of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to perform the steps of the dynamic processing method of microphone delay provided by any embodiment of the present application.
[0097] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0098] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the processor executes the steps of the dynamic processing method of microphone delay provided by any embodiment of the present application.
[0099] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0100] It should be noted that the terms used in this application are only for describing specific embodiments and are not intended to limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates an exception, the words "one", "an", "a kind of" and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, the elements defined by the sentence "comprise a..." do not exclude the presence of other identical elements in the process, method or device comprising the elements.
[0101] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the internal parts of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0102] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. The above is only the preferred implementation method of this application. It should be pointed out that due to the limitations of textual expression, there are objectively infinite specific structures. For ordinary technicians in this technical field, without departing from the principles of the present invention, they can also make several improvements, modifications or changes, and can also combine the above technical features in an appropriate manner; these improvements, modifications, changes or combinations, or the direct application of the inventive concept and technical solution to other occasions without improvement, should be regarded as the scope of protection of this application.
Claims
1. A method for dynamic processing of microphone delay, characterized in that: include: In response to a target audio algorithm calling a recording interface, determining a dynamic delay time adapted to the target audio algorithm; In response to the recording interface monitoring the sound signal, sending the sound signal other than the reference sound signal in the sound signal to the target audio algorithm, and waiting for the dynamic delay time before sending the reference sound signal to the target audio algorithm, so that the microphone delay time meets the requirements of the target algorithm; In which, the reference sound signal is the original digital signal corresponding to the sound signal played by the speaker, and the microphone delay time is the time difference between the first moment and the second moment. The first moment is the moment when the target audio algorithm receives the echo signal of the reference sound signal, and the second moment is the moment when the target audio algorithm receives the reference sound signal.
2. The method according to claim 1, characterized in that The determining of the dynamic delay time adapted to the target audio algorithm includes: Reading a dynamic delay time adapted to the target audio algorithm from a preset storage location; Alternatively, a reference delay time is determined according to a requirement of the target audio algorithm on microphone delay, and a dynamic delay time adapted to the target audio algorithm is determined according to the reference delay time.
3. The method according to claim 2, characterized in that The target audio algorithm's requirement for microphone delay includes a minimum delay time and a maximum delay time, and determining the reference delay time according to the target audio algorithm's requirement for microphone delay includes: Determining the minimum delay time as the reference delay time; Alternatively, the maximum delay time is determined as the reference delay time; Alternatively, the reference delay time is determined according to the minimum delay time and the maximum delay time.
4. The method according to claim 3, characterized in that Determining the reference delay time according to the minimum delay time and the maximum delay time includes: An average value of the minimum delay time and the maximum delay time is determined as the reference delay time.
5. The method according to claim 2, characterized in that The determining, according to the reference delay time, a dynamic delay time adapted to the target audio algorithm comprises: The difference between the tested delay time and the reference delay time is determined as the dynamic delay time.
6. The method according to claim 5, characterized in that Before calling the recording interface in response to the target audio algorithm and determining the dynamic delay time adapted to the target audio algorithm, the method further includes: Controlling the loudspeaker to play a preset sound signal and monitoring the sound signal through the recording interface at the same time, wherein the sound signal monitored by the recording interface includes a reference sound signal corresponding to the preset sound signal and an echo signal of the reference sound signal; The delay time of the test is determined according to the time when the preset audio algorithm receives the echo signal of the reference tone signal and the time when the reference tone signal is received.
7. The method according to claim 1, characterized in that Also includes: The target audio algorithm is used to eliminate the echo signal of the reference sound signal according to the reference sound signal.
8. A dynamic processing device for microphone delay, characterized in that: include: A determination module, configured to call a recording interface in response to a target audio algorithm and determine a dynamic delay time adapted to the target audio algorithm; a sending module, configured to send, in response to the recording interface monitoring a sound signal, a sound signal other than a reference sound signal in the sound signal to the target audio algorithm, and to wait for the dynamic delay time before sending the reference sound signal to the target audio algorithm, so that the microphone delay time meets the requirements of the target algorithm; In which, the reference sound signal is the original digital signal corresponding to the sound signal played by the speaker, and the microphone delay time is the time difference between the first moment and the second moment. The first moment is the moment when the target audio algorithm receives the echo signal of the reference sound signal, and the second moment is the moment when the target audio algorithm receives the reference sound signal.
9. An electronic device, characterized in that: The electronic device comprises: processor and memory; The processor is configured to execute the steps of the method for dynamic processing of microphone delay according to any one of claims 1 to 7 by calling the program or instructions stored in the memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute the steps of the method for dynamic processing of microphone delay according to any one of claims 1 to 7.
Citation Information
Patent Citations
Systems and methods for surround sound echo reduction
CN104429100A
Echo delay determination method and device, equipment and storage medium
CN113707160A
Echo cancellation method, voice recognition method, voice wake-up method and device
CN114203136A
Method for cancelling echo and related product
CN114360570A
Time delay calibration method for acoustic echo cancellation and television device
TWI736122B