Sound source positioning method, electronic device and storage medium

By performing noise reduction processing on the audio and fusing LiDAR point cloud information, the problem of sound source localization accuracy under the influence of reflective surfaces in the existing technology has been solved, and higher accuracy in determining the direction of the sound source has been achieved.

CN115762571BActive Publication Date: 2026-01-13AISPEECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211369438.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2026-01-13
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

Existing sound source localization methods have difficulty distinguishing between direct sound and reflected sound when there are reflecting surfaces, resulting in low localization accuracy. Furthermore, the noise from equipment operation interferes with the signal-to-noise ratio, affecting the algorithm's performance.

Method used

By denoising the audio, obtaining the speech angle spectrum, and fusing it with LiDAR point cloud information, the influence of reflective surfaces is shielded, and the final sound source location is output.

Benefits of technology

It improves the accuracy of sound source direction localization, reduces the interference of reflected sound on localization, and enhances the accuracy of sound source localization in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762571B_ABST
    Figure CN115762571B_ABST
Patent Text Reader

Abstract

The application discloses a sound source positioning method, an electronic device and a storage medium, wherein the method comprises: performing noise reduction processing on acquired first audio to obtain second audio, and calculating the azimuth spectrum information of the second audio in real time; judging whether the second audio matches a preset wake-up word; if the second audio matches the preset wake-up word, entering a command mode, using VAD to acquire the speech angle spectrum in the starting time of a speech segment in the second audio, then judging whether the first audio exists a strong reflection surface according to laser radar point cloud information; if the first audio exists the strong reflection surface, acquiring the speech angle spectrum in the starting and ending time of the speech segment in the second audio; based on the existence of the strong reflection surface, fusing the speech angle spectrum and the judgment of the laser radar point cloud information, and outputting the final sound source azimuth. The application fuses the speech angle spectrum of the audio after noise reduction processing and the judgment of the laser radar point cloud information, realizes the determination of the sound source azimuth, and can improve the accuracy of the sound source direction positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of sound source localization technology, and particularly relates to a sound source localization method, electronic device, and storage medium. Background Technology

[0002] The main method in existing technology is to select the output channel with the strongest energy from the wake-up channels based on beamforming during wake-up. The actual direction corresponding to this channel is used as the wake-up direction output. The main steps are: ① The device buffers a fixed duration of multi-channel audio after beamforming in real time, with each channel corresponding to a candidate direction. When the device is woken up, the buffered audio segment is used for direction estimation. ② Energy calculation is performed on some or all frequency bands of the beamformed audio, and the channel with the strongest energy is selected from the wake-up channels. The direction corresponding to the selected channel is the sound source direction.

[0003] Existing beamforming algorithms require significant computing power and memory, have low directional discrimination, and do not support a large number of alternative regions. Devices relying solely on sound information cannot distinguish between direct and reflected sound. When materials with high reflectivity (such as walls or mirrors) are present near the device, the time delay between reflected and direct sound is minimal, easily confusing the directions of arrival of direct and reflected sound. Reflected sound can affect the final location determination. During device operation, the noise from the motor rotor and friction between certain parts of the device and the ground can reduce the signal-to-noise ratio, causing the algorithm to fail. The accuracy of sound source direction localization is not high; generally, the number of alternative locations (regions) is similar to the number of microphones. It has poor resistance to interference and is easily affected by environmental noise. When the device is running, its own noise can interfere with the final result.

[0004] The inventors discovered that the existing sound source localization methods have low accuracy in locating the direction of the sound source, cannot distinguish between direct sound and reflected sound, and suffer from performance degradation during equipment operation, which reduces the signal-to-noise ratio and causes the algorithm to fail. Summary of the Invention

[0005] The embodiments of the present invention are intended to solve at least one of the above-mentioned technical problems.

[0006] In a first aspect, embodiments of the present invention provide a sound source localization method, comprising: performing noise reduction processing on an acquired first audio to obtain a second audio, and calculating the azimuth spectrum information of the second audio in real time; determining whether the second audio matches a preset wake-up word; if it matches the preset wake-up word, determining whether the first audio has a strong reflective surface based on lidar point cloud information; if it does, acquiring the speech angle spectrum within the start and end time of a segment in the second audio; if it does not exist, directly outputting the final sound source location; and based on the existence of a strong reflective surface, fusing the speech angle spectrum with the determination of the lidar point cloud information to output the final sound source location.

[0007] In a second aspect, embodiments of the present invention provide an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the sound source localization methods described above.

[0008] Thirdly, embodiments of the present invention provide a storage medium storing one or more programs including execution instructions, the execution instructions being readable and executed by electronic devices (including but not limited to computers, servers, or network devices, etc.) to perform any of the sound source localization methods described above.

[0009] Fourthly, embodiments of the present invention also provide a computer program product, the computer program product including a computer program stored on a storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform any of the above-described sound source localization methods.

[0010] This invention embodiment achieves the determination of the sound source location by performing noise reduction processing on audio with a strong reflective surface and obtaining the speech angle spectrum of the noise-reduced audio, which is then fused with the judgment of lidar point cloud information. At the same time, it can also improve the accuracy of sound source direction positioning. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of an embodiment of the sound source localization method of the present invention;

[0013] Figure 2 A flowchart of another embodiment of the sound source localization method of the present invention;

[0014] Figure 3 A flowchart illustrating the sound source localization process provided in an embodiment of the present invention;

[0015] Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device of the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0018] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, elements, data structures, etc., that perform a specific task or implement a specific abstract data type. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0019] In this invention, terms such as "module," "device," and "system" refer to relevant entities applied to a computer, such as hardware, combinations of hardware and software, software, or software in execution. More specifically, for example, an element can be, but is not limited to, a process running on a processor, a processor, an object, an executable element, an execution thread, a program, and / or a computer. Furthermore, an application program or script running on a server, and the server itself, can also be an element. One or more elements may be in an execution process and / or thread, and elements may be localized on a single computer and / or distributed across two or more computers, and may be run on various computer-readable media. Elements can also communicate via local and / or remote processes based on signals having one or more data packets, for example, signals from data interacting with another element in a local system, a distributed system, and / or interacting with other systems via signals over a network of the Internet.

[0020] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising" or "including" include not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0021] This invention provides a sound source localization method, which can be applied to electronic devices. The electronic device can be a computer, server, or other electronic product, etc., and this invention does not limit this to any particular device.

[0022] Please refer to Figure 1 This illustrates a sound source localization method provided by an embodiment of the present invention.

[0023] like Figure 1 As shown, in step 101, the acquired first audio is denoised to obtain the second audio, and the azimuth spectrum information of the second audio is calculated in real time.

[0024] In step 102, it is determined whether the second audio matches a preset wake-up word;

[0025] In step 103, if it matches the preset wake-up word, then it is determined whether the first audio has a strong reflective surface based on the lidar point cloud information;

[0026] In step 104, if the source exists, the speech angle spectrum within the start and end time of the second audio segment is obtained; if it does not exist, the final sound source location is directly output.

[0027] In step 105, based on the existence of a strong reflective surface, the speech angle spectrum is fused with the lidar point cloud information to output the final sound source location.

[0028] In this embodiment, for step 101, the beamforming (BF) module is used to perform noise reduction processing on the acquired first audio to obtain the noise-reduced second audio. The beamforming module can filter out residual noise energy and perform noise reduction on the sound in different directions according to the phase information to improve the signal-to-noise ratio. Then, the azimuth spectrum information of the noise-reduced second audio is calculated in real time by the direction of arrival (DOA) estimation module. The DOA estimation module can estimate the direction spectrum of the sound frame by frame in real time.

[0029] Next, in step 102, it is determined whether the noise-reduced second audio matches the preset wake-up word. The clean, noise-reduced audio is sent to the wake-up module to determine if it matches the wake-up word. Then, in step 103, if the noise-reduced second audio matches the preset wake-up word, it is determined whether there is a strong reflective surface in the first audio based on the LiDAR point cloud information. The LiDAR emits a signal, which is reflected when it encounters a reflective object and received by the LiDAR receiver. The LiDAR can obtain the three-dimensional spatial position of the reflection point based on information such as the delay, frequency shift, and rotation angle of the received and reflected signals, and output the result in the form of points. All points in space combined constitute the LiDAR point cloud information. For example, when the wake-up module determines that the input second audio matches the preset wake-up word, it determines whether there is a strong reflective surface in the direction of the current microphone of the device based on the LiDAR point cloud information.

[0030] Next, for step 104, if a strong reflective surface exists in the first audio, the speech angle spectrum within the start and end time of the segment in the denoised second audio is obtained. This is achieved through speech activation detection. The segment in the second audio is a command word segment, where a command word is a word used to issue a command to the device, such as saying "turn on" or "lower the volume" to a TV. The command word is not a wake-up word, but rather the voice signal following the wake-up word. If no strong reflective surface exists in the first audio, the final sound source location is directly output. The Doa algorithm searches for peaks in the spatial spectrum, where the x-axis is the candidate direction. When a reflective surface exists, the direction corresponding to the reflective surface on the x-axis is masked. When the maximum peak appears in the direction of the reflective surface, this peak is ignored, and a peak is searched in the non-reflective surface direction. When no reflective surface exists, no peaks are ignored.

[0031] Finally, for step 105, when there is a strong reflective surface in the first audio, the judgment of the speech angle spectrum and the lidar point cloud information are fused together, and the final sound source location is output. For example, the lidar point cloud information, the start and end time points of the speech segments in the second audio, and the audio azimuth spectrum information are fused together to finally output the sound source location information. Among them, the fusion of lidar point cloud information can shield the sound source mirror location, such as walls and mirrors, and avoid incorrect estimation of the reflection direction.

[0032] The method in this application embodiment performs noise reduction processing on audio with strong reflective surfaces and obtains the speech angle spectrum of the noise-reduced audio to fuse with the judgment of lidar point cloud information, thereby realizing the determination of the sound source location and improving the accuracy of sound source direction positioning.

[0033] In some optional embodiments, after the first audio is denoised to obtain the second audio, the azimuth spectrum of the sound in the second audio is estimated frame by frame in real time. The azimuth spectrum information of the second audio after denoising is calculated in real time using the angle of arrival estimation module, and the calculated azimuth spectrum information is output.

[0034] Please refer to Figure 2 This illustrates another sound source localization method provided by an embodiment of the present invention. The flowchart is mainly a flowchart... Figure 1 The flowchart of the step 102, "determining whether the second audio matches the preset wake word", is a further defined step.

[0035] like Figure 2 As shown, in step 201, if the second audio matches a preset wake-up word, the device enters command mode and performs voice activation detection on the second audio.

[0036] In step 202, if the second audio does not match the preset wake word, the second audio is determined to be invalid audio, and the acquisition is repeated.

[0037] In this embodiment, for step 201, if the noise-reduced second audio matches the preset wake-up word, the device enters command mode and performs voice activation detection on the noise-reduced second audio, using voice activation detection to determine the voice angle spectrum within the start and end time of the command word segment; then, for step 202, if the noise-reduced second audio does not match the preset wake-up word, the noise-reduced second audio is considered invalid audio, and the first audio is reacquired and noise-reduced again.

[0038] The method in this application embodiment performs voice activation detection on the second audio by determining that the second audio matches a preset wake word, thereby further improving the accuracy of sound localization in the process of sound localization.

[0039] In some optional embodiments, the system determines whether there is a strong reflective surface in the direction of the microphone device used to acquire the first audio based on the lidar point cloud information. When the wake-up module determines that the input voice matches the wake-up word, it determines whether there is a strong reflective surface in the direction of the microphone device based on the lidar point cloud information. The lidar point cloud information is used to assist in sound source localization and improve the performance of sound source localization in near-wall scenarios.

[0040] In some optional embodiments, based on a first audio segment with a strong reflective surface, speech activation detection is used to determine the start and end times of the segment corresponding to the preset wake-up word in the second audio, and the corresponding speech angle spectrum. Speech activation detection is also used to determine the start and end times of the command word segment corresponding to the preset wake-up word, and the corresponding speech angle spectrum within those times. Speech activation detection can determine the start time of the command word segment, avoiding the introduction of interference noise from noisy segments or the location information of background noise from silent segments. If no strong reflective surface exists, speech activation detection is used to determine the start and end times of the segment corresponding to the preset wake-up word in the second audio, and the corresponding speech angle spectrum. The segment corresponding to the preset wake-up word can be a command word segment following the wake-up word. If a strong reflective surface exists, the energy of the angle spectrum in the direction of the strong reflective surface is filtered out, and the speech direction is searched in the direction of the non-strong reflective surface.

[0041] It should be noted that this solution first uses the Adaptive Echo Cancelling (AEC) algorithm to initially eliminate local noise (including noise generated by local broadcasts and device movement). AEC is used to perform preliminary noise reduction on local playback and device movement noise. Then, the Beamforming (BF) module is used to further reduce noise on the output audio of AEC. Finally, the Direction of Arrival (DOA) estimation module is used to calculate the output azimuth spectrum information in real time. At the same time, the clean audio after BF is sent to the wake-up module to determine whether it is a wake-up word.

[0042] When the wake-up module determines that the input voice matches the wake-up word, it uses LiDAR point cloud information to determine whether there is a strong reflective surface in the direction of the device's microphone. Simultaneously, Voice Activity Detection (VAD) determines the speech angle spectrum within the start and end time of the command phrase. These results are then combined with the reflective surface location determination from the LiDAR point cloud to output the final sound source location.

[0043] In some optional embodiments, the lidar point cloud information, the start and end times of the segment corresponding to the preset wake-up word in the second audio, and the speech angle spectrum corresponding to the start and end times of the segment are fused together, and the final sound source location is output. For example, lidar information, the start and end times of the corresponding segment, and the speech angle spectrum information corresponding to the start and end times are fused together, and the sound source location information is finally output. For the fused lidar point cloud information, the sound source mirror location such as walls and mirrors can be masked, avoiding incorrect estimation of the reflection direction.

[0044] In some alternative embodiments, before performing noise reduction processing on the acquired first audio using beamforming, echo cancellation processing is also required. An adaptive echo cancellation algorithm is used to remove noise from the acquired audio, including local announcements and noise generated during device movement. After echo cancellation, the subsequent sound source localization can be more accurate. For example, AEC is used to perform preliminary noise reduction on local playback and device movement noise. The adaptive echo cancellation (AEC) algorithm can effectively perform preliminary noise reduction on local noise (including noise generated from local announcements and device movement).

[0045] It should be noted that this application uses a microphone device to acquire the first audio and performs a first noise reduction process on the first audio. The first noise reduction process removes noise from the first audio acquired by the microphone, including noise generated by the local broadcast and the movement of the device. After the first noise reduction process, a second noise reduction process is required. The second noise reduction process uses a beamforming (BF) module to further reduce the noise of the audio after the first noise reduction. The beamforming (BF) module can effectively process external interference. The first audio is acquired through a microphone device, which consists of a microphone array and can acquire more accurate audio information.

[0046] It should be noted that this application also provides an alternative solution: instead of using the VAD module, DOA estimation is performed using audio within a fixed frame length range after wake-up. This solution is simple to implement, but it limits the interval between the command word and the wake-up word, affecting the user experience. Not using AEC and / or BF noise reduction modules can reduce computational requirements and simplify algorithm complexity. However, it can only be used in quiet environments. Determining the sound source location solely using the speech signal without using external information is also an option, but the accuracy of location determination decreases in environments with strong reflections, such as near walls or corners. In this application, algorithms such as Inter-channel Phase Difference (IPD), Generalized Cross Correlation (GCC), and Multiple Signal Classification (MUSIC) can be used to replace the compressed sensing-based DOA algorithm module. Alternatively, other noise reduction algorithms or neural network algorithms can be used to replace the BF module for noise reduction processing. This application does not limit the use of image or video information to replace the point cloud information of lidar in determining the location of sound sources such as walls and mirrors.

[0047] Please refer to Figure 3 The document presents a flowchart illustrating the implementation process of the sound source localization method of the present invention.

[0048] like Figure 3As shown, Step 1: Use AEC to perform preliminary noise reduction on the playback and motion noise of the device.

[0049] Step 2: Use the BF module to further reduce external noise interference.

[0050] Step 3: Use the DOA module to estimate the orientation spectrum of the denoised audio in real time.

[0051] Step 4: Use the wake-up model to determine if it is a wake word.

[0052] Step 5: If it is a wake word, use VAD to determine the start and end times of the command word in the audio following the wake word.

[0053] Step 6: Fuse the lidar signal, the start and end times of the command words, and the audio azimuth spectrum information to finally output the azimuth information of the sound source.

[0054] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0055] In some embodiments, the present invention provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions, which can be read and executed by electronic devices (including but not limited to computers, servers, or network devices) to perform any of the sound source localization methods described above.

[0056] In some embodiments, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform any of the above-described sound source localization methods.

[0057] In some embodiments, the present invention also provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a sound source localization method.

[0058] Figure 4 This is a schematic diagram of the hardware structure of an electronic device for performing a sound source localization method according to another embodiment of this application, as shown below. Figure 4 As shown, the device includes:

[0059] One or more processors 410 and memory 420, Figure 4 Take a processor 410 as an example.

[0060] The device for performing the sound source localization method may further include an input device 430 and an output device 440.

[0061] The processor 410, memory 420, input device 430, and output device 440 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.

[0062] The memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the sound source localization method in the embodiments of this application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 420, thereby implementing the sound source localization method in the above-described method embodiments.

[0063] The memory 420 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the sound source localization device. Furthermore, the memory 420 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 420 may optionally include memory remotely located relative to the processor 410, and this remote memory may be connected to the sound source localization device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0064] Input device 430 can receive input digital or character information, and generate signals related to user settings and function control of the sound source localization device. Output device 440 may include display devices such as a display screen.

[0065] The one or more modules are stored in the memory 420, and when executed by the one or more processors 410, they execute the sound source localization method in any of the above method embodiments.

[0066] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0067] The electronic devices in this application embodiments exist in various forms, including but not limited to:

[0068] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.

[0069] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include: PDAs, MIDs, and UMPCs, etc.

[0070] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes audio and video players, handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.

[0071] (4) Other airborne electronic devices with data interaction capabilities, such as vehicle-mounted systems installed on vehicles.

[0072] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0073] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for locating a sound source, comprising: performing noise reduction on a first audio to obtain a second audio by using beamforming, and calculating a bearing spectrum of the second audio in real time; determining whether the second audio matches a preset wake-up word; if the second audio matches the preset wake-up word, determining whether there is a strong reflection surface for the first audio according to laser radar point cloud information; if there is a strong reflection surface, obtaining a speech bearing spectrum of a speech segment in a start and end time of the second audio, and if there is not a strong reflection surface, directly outputting a final sound source bearing; based on the strong reflection surface, fusing the speech bearing spectrum and the determination of the laser radar point cloud information to output the final sound source bearing.

2. The method of claim 1, wherein, the calculating the bearing spectrum of the second audio in real time comprises: estimating a bearing spectrum of a sound in the second audio in real time frame by frame.

3. The method of claim 1, wherein, the determining whether the second audio matches the preset wake-up word comprises: if the second audio matches the preset wake-up word, entering a command mode of a device, and performing speech activation detection on the second audio; if the second audio does not match the preset wake-up word, determining that the second audio is invalid audio, and re-performing collection.

4. The method of claim 3, wherein, the determining whether there is a strong reflection surface for the first audio according to the laser radar point cloud information if the second audio matches the preset wake-up word comprises: determining whether there is a strong reflection surface in a direction where a microphone device for collecting the first audio is located based on the laser radar point cloud information.

5. The method of claim 4, wherein, the method further comprises: if there is a strong reflection surface, determining a start and end time of a speech segment corresponding to the preset wake-up word in the second audio and a speech bearing spectrum corresponding to the start and end time of the speech segment by using the speech activation detection.

6. The method of claim 1, wherein, the fusing the speech bearing spectrum and the determination of the laser radar point cloud information to output the final sound source bearing comprises: fusing the laser radar point cloud information, the start and end time of the speech segment corresponding to the preset wake-up word in the second audio, and the speech bearing spectrum corresponding to the start and end time of the speech segment to output the final sound source bearing.

7. The method of claim 1, wherein, the method further comprises: before the performing noise reduction on the first audio to obtain the second audio by using beamforming, removing noise containing local broadcasting and noise generated when a device moves by using an adaptive echo cancellation algorithm.

8. The method of claim 7, wherein, the method further comprises: collecting the first audio by using a microphone device, and performing echo cancellation on the first audio, wherein the microphone device is a microphone array.

9. An electronic device comprising: at least one processor, and a memory connected to the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method of any one of claims 1 to 8.

10. A storage medium having stored thereon a computer program, characterized in that the program is executed by a processor to implement the steps of the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Curvelet domain statistics self-adaptive threshold ground penetrating radar data de-noising method and system

    CN109581516A

  • Voice control method and device for service robot

    CN112562671A