Sound source positioning method and device, equipment and storage medium
By combining phase difference and energy distribution data for sound source localization, the problem of inaccurate differentiation between in-vehicle and out-of-vehicle voice signals in automotive voice interaction systems has been solved, achieving higher accuracy and safety.
Patent Information
- Application Number
- CN202511754707.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-01-20
AI Technical Summary
Existing car voice interaction systems cannot accurately distinguish between voice signals inside and outside the vehicle, leading to external voice signals being misidentified as internal voice responses, which affects the safety of the interaction.
By acquiring the phase difference and energy distribution data of the target sound source signal arriving at different microphones in the microphone array, and combining these data, the sound source is located to determine whether the sound source signal originates from inside or outside the vehicle.
It improves the accuracy of sound source signal differentiation, reduces misjudgments caused by noise and other factors, and ensures the safety and accurate response of the voice interaction system.
Smart Images

Figure CN121364446A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a sound source positioning method and device, equipment and storage medium. BACKGROUND
[0002] With the development of intelligent vehicles, voice interaction has become one of the core ways of vehicle human-computer interaction. The current vehicle voice interaction system cannot accurately distinguish whether the voice signal is from inside or outside the vehicle, so it is easy to misrecognize the voice outside the vehicle as the voice inside the vehicle for response, affecting the safety of interaction. SUMMARY
[0003] The technical problem solved by the present application is to provide a sound source positioning method, device and storage medium, which can accurately distinguish the sound source signals inside and outside the vehicle.
[0004] To solve the above technical problems, one technical solution adopted by the present application is to provide a sound source positioning method, which comprises: obtaining the phase difference of a target sound source signal reaching different microphones in a microphone array as a target phase difference, and obtaining energy distribution data of the target sound source signal; and comprehensively positioning the sound source based on the target phase difference and the energy distribution data to determine a target positioning result of the target sound source signal; the target positioning result is used to represent whether the target sound source signal is from inside or outside the vehicle.
[0005] To solve the above technical problems, another technical solution adopted by the present application is to provide a sound source positioning device, which comprises: an acquisition module and a determination module; the acquisition module is used to obtain the phase difference of a target sound source signal reaching different microphones in a microphone array as a target phase difference, and obtain energy distribution data of the target sound source signal; and the determination module is used to comprehensively position the sound source based on the target phase difference and the energy distribution data to determine a target positioning result of the target sound source signal; the target positioning result is used to represent whether the target sound source signal is from inside or outside the vehicle.
[0006] To solve the above technical problems, still another technical solution adopted by the present application is to provide an electronic device, which comprises a storage and a processor coupled with each other, and the storage stores program instructions; the processor is used to execute the program instructions stored in the storage to realize the above method.
[0007] To solve the above technical problems, yet another technical solution adopted by the present application is to provide a computer readable storage medium for storing program instructions, which can be executed to realize the above method.
[0008] The above scheme combines the phase difference of the target sound source signal reaching different microphones in the microphone array and the energy distribution data of the target sound source signal to perform sound source positioning, and determines the positioning result of the target sound source signal. The phase difference can reflect the spatial position of the sound source relative to the microphone array, and the energy distribution can reflect the propagation characteristics of the signal. Therefore, the sound source positioning method combining the phase difference and the energy distribution data of the sound signal can combine the spatial characteristics and the propagation characteristics of the signal for positioning. Compared with the method of using only the energy size relationship to determine the sound source inside or outside the vehicle, the signal information used by the present application is more comprehensive, so that the misjudgment caused by noise and the like when only using energy comparison can be effectively reduced, and the sound source signal inside or outside the vehicle can be accurately distinguished. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 FIG. 1 is a flow diagram of an embodiment of a sound source positioning method provided by the present application; Figure 2 FIG. 2 is a flow diagram of an embodiment of step S12 shown in FIG. 1; Figure 1 Figure 3 FIG. 3 is a schematic diagram of an embodiment of a sound source positioning device provided by the present application; Figure 4 FIG. 4 is a schematic diagram of an embodiment of an electronic device provided by the present application; Figure 5 FIG. 5 is a schematic diagram of a computer readable storage medium provided by the present application. DETAILED DESCRIPTION
[0010] In order to make the purpose, technical scheme and effect of the present application more clear and explicit, the present application will be further described in detail below with reference to the drawings and embodiments.
[0011] In addition, if the present application embodiments involve descriptions such as "first", "second", etc., the descriptions of "first", "second", etc. are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first" and "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the realization of a person skilled in the art, and when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, and is not within the protection scope required by the present application.
[0012] It should be noted that the method provided in the present application can be used in any scenario where it is necessary to accurately distinguish between in-vehicle and out-of-vehicle sound source signals, and can but is not limited to a voice interaction scenario. For example, in a voice interaction scenario, a voice assistant is provided in the vehicle, and in order to avoid the voice assistant being mistakenly woken up by an out-of-vehicle voice signal or the voice assistant mistakenly responding to an out-of-vehicle command, it is necessary to accurately distinguish between in-vehicle and out-of-vehicle sound source signals, i.e., to accurately distinguish whether the currently received sound source signal is from outside the vehicle.
[0013] For ease of description, the following is described in the form of a voice interaction scenario.
[0014] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the sound source positioning method provided in the present application. It should be noted that the present embodiment is not limited to the flow sequence shown in Figure 1 . As shown in Figure 1 , the present embodiment includes: S11: Obtain the phase difference of a target sound source signal arriving at different microphones in a microphone array as a target phase difference, and obtain energy distribution data of the target sound source signal.
[0015] In the present embodiment, the microphone array is provided in the vehicle, and in addition, in order to make the phase difference of the same voice signal arriving at different microphones in the microphone array more obvious, the distance between different microphones in the microphone array can be greater than a preset distance threshold when the microphone array is provided; the preset distance threshold can be set as needed.
[0016] In a specific embodiment, the preset distance threshold is 60 centimeters, i.e., the distance between different microphones in the microphone array is greater than 60 centimeters.
[0017] In an implementation scenario, the target sound source signal is the original sound source signal currently received.
[0018] In another implementation scenario, in order to avoid the voice assistant being mistakenly woken up or responding, and to avoid sound source distinguishing processing of signals irrelevant to the wake-up and response, resulting in waste of processing resources, after the original sound source signal is obtained, it can be detected whether there is a wake-up word or a command word in the original sound source signal, and if no such word is detected, the sound source positioning method of the present application does not need to be performed.
[0019] But if the presence of the wake-up word or the command word is detected, the sound source signal in the original sound source signal related to the wake-up word or the command word can be taken as the target sound source signal; or based on the wake-up word or the command word detected in the original sound source signal, the wake-up time (the pronunciation time of the voice signal of the wake-up word or the command word) is determined; and the signal in the original sound source signal within a preset time range before and after the wake-up time is extracted as the target sound source signal. The signal in the original sound source signal within a preset time range before and after the wake-up time is taken as the target sound source signal, and the target sound source signal determined in this way includes the voice signal within the first preset time range before the wake-up time and the second preset time range after the wake-up time. The first preset time range and the second preset time range can be the same or different, and the voice signal of a certain length can be extracted as the target sound source signal according to actual needs.
[0020] In a specific embodiment, the first preset time range is 500ms-1s, and the second preset time range is 1s-2s.
[0021] In some embodiments, after obtaining the original sound source signal, the original sound source signal can also be subjected to noise reduction and gain calibration processing, and the clean voice signal obtained after processing is taken as a new original sound source signal; and then the above-mentioned scheme is used to obtain the target sound source signal based on the original sound source signal.
[0022] In an implementation, after obtaining the target sound source signal, the time difference (TDOA) of the target sound source signal arriving at different microphones in the microphone array can be calculated, and the time difference is converted into a phase difference (phase difference = 2π x distance difference / wavelength, distance difference = time difference x sound wave speed) according to the sound wave propagation speed (340m / s).
[0023] The following explains how to obtain the energy distribution data of the target sound source signal.
[0024] In this embodiment, the energy distribution data of the target sound source signal includes the energy proportion of the target sound source signal in each sub-band, and the sub-band is obtained by frequency band division on the frequency range of the target sound source signal, and the energy proportion of the sub-band is the ratio of the energy of the corresponding sub-band to the total energy of the frequency range.
[0025] For example, the frequency range of the target sound source signal is 20Hz-20kHz, and the frequency range is divided into three frequency bands: a low frequency band 20Hz-300Hz, a medium frequency band 300Hz-3kHz, and a high frequency band 3kHz-20kHz.
[0026] Among them, the low frequency band (20Hz-300Hz) mainly corresponds to the frequency band of the environmental noise such as the wind noise outside the vehicle and the engine noise; the medium frequency band (300Hz-3kHz) is mainly the core frequency band of the human voice; the high frequency band (3kHz-20kHz): the attenuation amplitude of the environmental noise outside the vehicle in the frequency band is significantly greater than that in the vehicle, and the high frequency component of the voice in the vehicle is more complete.
[0027] In an embodiment, the effective speech frame in the target sound source signal can be Fourier transformed, and the energy proportion of each sub-band (energy of sub-band / total energy) is calculated.
[0028] S12: The target sound source signal is sent to the voice interaction module of the vehicle for subsequent operation if the target positioning result indicates that the target sound source signal is from inside the vehicle, and no response is made to the target sound source signal if the target positioning result indicates that the target sound source signal is from outside the vehicle, thereby realizing false wake-up or false response.
[0029] In an implementation scenario, if the target positioning result indicates that the target sound source signal is from inside the vehicle, the target sound source signal is sent to the voice interaction module of the vehicle for subsequent operation; if the target positioning result indicates that the target sound source signal is from outside the vehicle, no response is made to the target sound source signal, thereby realizing false wake-up or false response.
[0030] The following briefly introduces how to perform sound source positioning by comprehensively using the target phase difference and the energy distribution data to determine the target positioning result of the target sound source signal.
[0031] In an embodiment, please refer to Figure 2 , Figure 2 is Figure 1 a flowchart of an embodiment of step S12 shown in FIG. 2. In this embodiment, step S12 further includes: S21: Obtain the matching result between the target phase difference and each reference phase difference; the matching result is used to indicate whether there is a matching phase difference in the target phase difference among the reference phase differences.
[0032] In this embodiment, the microphone array is arranged in the vehicle, the reference phase difference is the phase difference of the reference sound source signal from inside the vehicle to different microphones in the microphone array, and the different reference sound source signals are the sound source signals from different preset positions inside the vehicle.
[0033] For example, the microphone array is composed of at least 4 microphones distributed at different positions in the vehicle (such as the driver's side A pillar, the co-pilot's side A pillar, the rear seat roof armrest, etc.), which is used to synchronously collect the voice signals inside and outside the vehicle.
[0034] The reference sound source signals are sound source signals from different preset positions in the vehicle (for example, the driver's seat, the front passenger seat, and the seats in the back row); the reference phase difference is a phase difference between the reference sound source signals from the vehicle and different microphones in the microphone array; wherein, for each preset position, the corresponding reference phase difference can be determined by using a plurality of sets of reference sound source signals.
[0035] In an embodiment, the similarity between the target phase difference and each reference phase difference can be obtained, and it is determined whether the maximum similarity is greater than a preset similarity threshold; if the maximum similarity is greater than the preset similarity threshold (for example, 0.8), the reference phase difference corresponding to the maximum similarity threshold is taken as the matching phase difference matched with the target phase difference; if the maximum similarity is not greater than the preset similarity threshold, it is determined that there is no matching phase difference matched with the target phase difference in each reference phase difference.
[0036] It can be understood that if there is a matching phase difference, it can be preliminarily determined that the target sound source signal comes from inside the vehicle.
[0037] S22: Based on the matching result and the energy distribution data, the sound source positioning is performed to determine a target positioning result.
[0038] In this embodiment, after the matching result between the target phase difference and each reference phase difference is determined, the corresponding target positioning result can be determined based on the matching result and the energy proportion of the target sound source signal in at least one sub-frequency band obtained by division. The at least one sub-frequency band includes a first frequency band, and the first frequency band includes a frequency band corresponding to human voice. As described above, the first frequency band must contain the above-mentioned medium frequency band 300Hz-3kHz.
[0039] In a specific embodiment, based on the matching result and the energy proportion of the target sound source signal in at least one sub-frequency band obtained by division, the corresponding target positioning result is determined, including the following steps: First, based on the matching result between the target phase difference and each reference phase difference, and the energy proportion of the target sound source signal in the first frequency band, a preliminary positioning result of the target sound source signal is determined.
[0040] In this embodiment, if the matching result between the target phase difference and each reference phase difference indicates that there is a matching phase difference matched with the target phase difference, it is further detected whether the energy proportion of the first frequency band satisfies an energy threshold in the vehicle. If the energy proportion of the first frequency band satisfies the energy threshold in the vehicle, it is determined that the preliminary positioning result is that the target sound source signal comes from inside the vehicle and is from a preset position where the reference sound source signal corresponding to the matching phase difference is located; otherwise, it is determined that the preliminary positioning result indicates that the target sound source signal comes from outside the vehicle.
[0041] If the matching result between the target phase difference and each reference phase difference is that there is no matching phase difference matching the target phase difference, it is determined that the preliminary positioning result is that the target sound source signal is from outside the vehicle.
[0042] It should be noted that the energy threshold in the vehicle is an energy proportion threshold capable of distinguishing whether the sound source signal is in the vehicle or outside the vehicle, which is determined by a large number of real vehicle tests (a large number of in-vehicle sound source signals and a large number of out-of-vehicle sound source signals). The energy threshold is the energy proportion threshold of the sound source signal in the first frequency band.
[0043] In an embodiment, the first frequency band only includes the above-mentioned medium frequency band, and in other embodiments, the first frequency band includes both the above-mentioned medium frequency band and the above-mentioned low frequency band.
[0044] In a specific implementation scenario, the first frequency band includes the above-mentioned medium frequency band and low frequency band; and the above-mentioned energy proportion threshold includes a medium frequency band energy proportion threshold and a low frequency band energy proportion threshold. For example, the lower limit of the medium frequency band energy proportion threshold is 60%, and the upper limit of the low frequency band energy proportion threshold is 25%. That is, when the medium frequency band energy proportion of the target sound source signal is ≥ 60% and the low frequency band energy proportion is ≤ 25%, it is considered that the energy proportion of the first frequency band meets the energy threshold in the vehicle, and the preliminary positioning result is that the target sound source signal is from the vehicle and is in the preset position of the reference sound source signal corresponding to the matching phase difference. However, if the medium frequency band energy proportion of the target sound source signal is > 60% or the low frequency band energy proportion is > 25%, it is considered that the energy proportion of the first frequency band does not meet the energy threshold in the vehicle, and it is determined that the preliminary positioning result represents that the target sound source signal is from outside the vehicle.
[0045] Second, based on the preliminary positioning result, a corresponding target positioning result is determined.
[0046] In this embodiment, if the preliminary positioning result represents that the target sound source signal is from the vehicle, the target positioning result is determined to be that the target sound source signal is from the vehicle; and / or, if the preliminary positioning result represents that the target sound source signal is from outside the vehicle, the preliminary positioning result is verified by using the energy proportion of the target sound source signal in the second frequency band to obtain a corresponding verification result; and the target positioning result is determined by using the verification result, wherein the second frequency band is one of the at least one sub-frequency band, and the frequency of the second frequency band is higher than that of the first frequency band.
[0047] Here, the second frequency band can be understood as the high frequency band described above. That is, in this embodiment, if the preliminary positioning result represents that the target sound source signal is from outside the vehicle, the preliminary positioning result is verified by using the energy proportion in the high frequency band to obtain a corresponding verification result; and then the target positioning result is determined by using the verification result.
[0048] In an embodiment, if the energy proportion of the second frequency band is not greater than a preset energy threshold, it is determined that the verification result is that the preliminary positioning result is correct; otherwise, it is determined that the verification result is that the preliminary positioning result is incorrect.
[0049] If the verification result is that the preliminary positioning result is correct, it is determined that the target positioning result is that the target sound source signal is from outside the vehicle; if the verification result is that the preliminary positioning result is incorrect, it is determined that the target positioning result is that the target sound source signal is from inside the vehicle.
[0050] It should be noted that the preset energy threshold in the present embodiment is an energy proportion threshold capable of distinguishing between sound source signals inside and outside the vehicle, which is determined through a large number of real vehicle tests. The preset energy threshold is an energy proportion threshold of the sound source signal in the second frequency band. For example, if the energy proportion of the second frequency band is ≤15%, it is determined that the verification result is that the preliminary positioning result is correct, and the target positioning result is that the target sound source signal is from outside the vehicle; otherwise, it is determined that the verification result is that the preliminary positioning result is incorrect, and the target positioning result is that the target sound source signal is from inside the vehicle.
[0051] It should be noted that if the verification result is determined to be that the preliminary positioning result is incorrect and the target positioning result is that the target sound source signal is from inside the vehicle through the energy proportion verification of the second frequency band, it is determined that the energy proportion in the second frequency band is caused by sudden noise (such as engine noise) inside the vehicle. Therefore, when it is determined that the verification result is that the preliminary positioning result is incorrect, a related reminder can be sent to the driver.
[0052] The above scheme combines the phase difference of the target sound source signal arriving at different microphones in the microphone array and the energy distribution data of the target sound source signal to perform sound source positioning and determine the positioning result of the target sound source signal. The phase difference can reflect the spatial position of the sound source relative to the microphone array, and the energy distribution can reflect the propagation characteristics of the signal. Therefore, the sound source positioning method combining the phase difference and the energy distribution data of the sound signal can combine the spatial characteristics and the propagation characteristics of the signal to perform positioning. Compared with the method of determining the sound source inside and outside the vehicle by only using the energy size relationship, the signal information used in the present application is more comprehensive, so that the misjudgment caused by noise and the like when only using energy comparison can be effectively reduced, and the accuracy of distinguishing the sound source signal inside and outside the vehicle can be improved.
[0053] In a specific implementation scenario, the microphone array is arranged in the vehicle and is composed of 4 microphones distributed at different positions in the vehicle (such as the driver's side A pillar, the co-driver's side A pillar, and the rear seat overhead armrest), for synchronously collecting the voice signals (the original sound source signals in the above) inside and outside the vehicle; then the original signals collected by the microphone array are subjected to noise reduction (such as adaptive filter noise reduction), gain calibration processing, to eliminate the influence of environmental noise and microphone hardware differences; based on the preprocessed signals, the time difference of the signals reaching different microphones is calculated, and the phase difference is converted according to the sound wave propagation speed, as the target phase difference, and based on the wake-up word (such as "Hello, XX", or "open the door" command word) in the preprocessed signal, the wake-up moment (the moment corresponding to the wake-up word) is determined, and the pre-signal in the first time range before the wake-up moment and the post-signal in the second time range after the wake-up moment are extracted, to form the target sound source signal.
[0054] Then, it is judged whether there is a matching phase difference in the reference phase differences that matches the target phase difference, if not, it is preliminarily determined that the target sound source signal is from outside the vehicle; if there is, the energy ratio of the target sound source signal in the first frequency band is analyzed, if it meets the in-vehicle energy threshold, it is determined that the position of the target sound source signal is in the vehicle, and is the preset position corresponding to the matching reference phase difference (such as the driver's seat in the vehicle); if it does not meet, it is preliminarily determined that the target sound source signal is from outside the vehicle.
[0055] Among them, if it is preliminarily determined that the target sound source signal is from outside the vehicle, the energy ratio of the second frequency band is further used to verify whether the preliminary determination result is correct, wherein if the energy ratio of the second frequency band is less than or equal to the preset energy threshold, it is determined that the preliminary determination result is correct, and the target sound source signal is from outside the vehicle; otherwise, it is determined that the target sound source signal is from inside the vehicle.
[0056] Among them, if it is determined that the target sound source signal is from inside the vehicle, the target sound source signal is sent to the voice interaction module of the vehicle, and if it is from outside the vehicle, it is not sent to the voice interaction module.
[0057] It should be noted that the above scheme can accurately distinguish the sound source signals inside and outside the vehicle even in the wake-up word-free scene and the micro-window opening scene, and can meet the position judgment demand of the seat level in the vehicle.
[0058] In an embodiment, please refer to Figure 3 , Figure 3is a frame diagram of an embodiment of the sound source positioning apparatus provided in the present application. In the embodiment, the sound source positioning apparatus 30 comprises an acquisition module 31 and a determination module 32. The acquisition module 31 is configured to acquire a phase difference of a target sound source signal reaching different microphones in a microphone array as a target phase difference, and acquire energy distribution data of the target sound source signal; and the determination module 32 is configured to perform sound source positioning by comprehensively considering the target phase difference and the energy distribution data, and determine a target positioning result of the target sound source signal; the target positioning result is used to represent that the target sound source signal is from inside or outside the vehicle.
[0059] In some embodiments, the microphone array is arranged in the vehicle, and the distance between different microphones is greater than a preset distance threshold.
[0060] In some embodiments, the determination module 32 performs sound source positioning by comprehensively considering the target phase difference and the energy distribution data, and determines the target positioning result of the target sound source signal, including: acquiring a matching result between the target phase difference and each reference phase difference; the matching result is used to represent whether there is a matching phase difference in the reference phase difference that matches the target phase difference; and performing sound source positioning based on the matching result and the energy distribution data to determine the target positioning result.
[0061] In some embodiments, the energy distribution data of the target sound source signal includes: an energy proportion of the target sound source signal in each sub-frequency band, the sub-frequency band being obtained by frequency band division on a frequency range of the target sound source signal, and the energy proportion of the sub-frequency band being a ratio of energy of the corresponding sub-frequency band to total energy of the frequency range; and performing sound source positioning based on the matching result and the energy distribution data to determine the target positioning result, including: determining the target positioning result based on the matching result and the energy proportion of the target sound source signal in at least one sub-frequency band.
[0062] In some embodiments, the at least one sub-frequency band includes a first frequency band, and the first frequency band includes a frequency band corresponding to human voice; and determining the target positioning result based on the matching result and the energy proportion of the target sound source signal in at least one sub-frequency band, including: determining a preliminary positioning result of the target sound source signal based on the matching result and the energy proportion of the target sound source signal in the first frequency band corresponding to the first frequency band; and determining the target positioning result based on the preliminary positioning result.
[0063] In some embodiments, the microphone array is arranged in the vehicle, the reference phase difference is a phase difference of reference sound source signals from different preset positions in the vehicle to different microphones in the microphone array, and the different reference sound source signals are sound source signals from different preset positions in the vehicle; the preliminary positioning result of the target sound source signal is determined based on the matching result and an energy proportion of the target sound source signal in the first frequency band, including: in response to the matching result representing that there is a matching phase difference, detecting whether the energy proportion of the first frequency band meets an energy threshold in the vehicle; if the energy proportion of the first frequency band meets the energy threshold in the vehicle, determining that the preliminary positioning result is that the target sound source signal is from the vehicle and is from a preset position corresponding to the reference sound source signal of the matching phase difference; otherwise, determining that the preliminary positioning result represents that the target sound source signal is from outside the vehicle; and in response to the matching result being that there is no matching phase difference, determining that the preliminary positioning result is that the target sound source signal is from outside the vehicle.
[0064] In some embodiments, the target positioning result is determined based on the preliminary positioning result, including: in response to the preliminary positioning result representing that the target sound source signal is from the vehicle, determining that the target positioning result is that the target sound source signal is from the vehicle; and / or, in response to the preliminary positioning result representing that the target sound source signal is from outside the vehicle, verifying the preliminary positioning result by using an energy proportion of the target sound source signal in a second frequency band to obtain a corresponding verification result; and determining the target positioning result by using the verification result, wherein the second frequency band is one of the at least one sub-frequency band, and the frequency of the second frequency band is higher than that of the first frequency band.
[0065] In some embodiments, the verification result is obtained by verifying the preliminary positioning result by using the energy proportion of the target sound source signal in the second frequency band, including: in response to the energy proportion of the second frequency band being not greater than a preset energy threshold, determining that the verification result is that the preliminary positioning result is correct; otherwise, determining that the verification result is that the preliminary positioning result is incorrect; and the target positioning result is determined by using the verification result, including: in response to the verification result being that the preliminary positioning result is correct, determining that the target positioning result is that the target sound source signal is from outside the vehicle; and in response to the verification result being that the preliminary positioning result is incorrect, determining that the target positioning result is that the target sound source signal is from the vehicle.
[0066] In some embodiments, after the target positioning result of the target sound source signal is determined, the sound source positioning device is further configured to: in response to the target positioning result representing that the target sound source signal is from the vehicle, sending the target sound source signal to a voice interaction module of the vehicle; in response to the target positioning result representing that the target sound source signal is from outside the vehicle, not responding to the target sound source signal; and obtaining the target sound source signal, including: obtaining an original sound source signal; determining a wake-up time based on a wake-up word or a command word in the detected original sound source signal; and extracting a signal in the original sound source signal within a preset time range before and after the wake-up time as the target sound source signal.
[0067] Please refer toFigure 4 , Figure 4 is a framework schematic diagram of an embodiment of an electronic device provided by the present application. In the embodiment, the electronic device 40 comprises a memory 41 and a processor 42 coupled with each other.
[0068] The memory 41 stores program instructions, and the processor 42 is configured to execute the program instructions stored in the memory 41 to implement the steps of any of the above methods. In a specific implementation scenario, the electronic device 40 can include but is not limited to a microcomputer, a server, and in addition, the electronic device 40 can also include a notebook computer, a tablet computer and other mobile devices, which are not limited herein.
[0069] Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps of any of the above embodiments. The processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 can be an integrated circuit chip having a processing capability of signals. The processor 42 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. In addition, the processor 42 can be jointly implemented by integrated circuit chips.
[0070] Please refer to Figure 5 , Figure 5 is a framework schematic diagram of a computer readable storage medium provided by the present application. The computer readable storage medium 50 of the embodiment of the present application stores program instructions 51, which when executed implement the method provided by any of the above embodiments and any non-conflicting combination. Wherein the program instructions 51 can form a program file and be stored in the above computer readable storage medium 50 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) executes all or part of the steps of the methods of various embodiments of the present application. And the aforementioned computer readable storage medium 50 includes: a U disk, a mobile hard disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk or an optical disk and various media that can store program codes, or a computer, a server, a mobile phone, a tablet and other terminal devices.
[0071] The above scheme combines the phase difference of the target sound source signal reaching different microphones in the microphone array and the energy distribution data of the target sound source signal to perform sound source positioning, and determines the positioning result of the target sound source signal. The phase difference can reflect the spatial position of the sound source relative to the microphone array, and the energy distribution can reflect the propagation characteristics of the signal. Therefore, the sound source positioning method combining the phase difference and the energy distribution data of the sound signal can combine the spatial characteristics and the propagation characteristics of the signal to perform positioning. Compared with the method of using only the energy size relationship to determine the sound source inside or outside the vehicle, the signal information used in the present application is more comprehensive, so that the misjudgment caused by noise and the like when only the energy is compared can be effectively reduced, and the accuracy of distinguishing the sound source signal inside or outside the vehicle can be improved.
[0072] In some embodiments, the device provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, details are not repeated here.
[0073] The above description of various embodiments tends to emphasize the differences between various embodiments, and the same or similar parts can be mutually referred to. For brevity, details are not repeated here.
[0074] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic, for example, the division of modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual units can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0075] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.
[0076] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0077] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0078] The above description is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A method of acoustic source localization, the method comprising: The method comprises: obtaining a phase difference of a target sound source signal reaching different microphones in a microphone array as a target phase difference, and obtaining energy distribution data of the target sound source signal; combining the target phase difference and the energy distribution data to perform sound source positioning to determine a target positioning result of the target sound source signal; the target positioning result is used to represent that the target sound source signal is derived from inside or outside the vehicle.
2. The method of claim 1, wherein, The microphone array is arranged in the vehicle, and the distance between different microphones is greater than a preset distance threshold.
3. The method of claim 1, wherein, The combining the target phase difference and the energy distribution data to perform sound source positioning to determine the target positioning result of the target sound source signal comprises: obtaining a matching result between the target phase difference and each reference phase difference; the matching result is used to represent whether there is a matching phase difference matching the target phase difference in each reference phase difference; based on the matching result and the energy distribution data, performing sound source positioning to determine the target positioning result.
4. The method of claim 3, wherein, The energy distribution data of the target sound source signal comprises: energy proportions of the target sound source signal in each sub-band, the sub-band being obtained by frequency band division on a frequency range of the target sound source signal, and the energy proportion of the sub-band being a ratio of energy of the corresponding sub-band to total energy of the frequency range; The sound source positioning based on the matching result and the energy distribution data to determine the target positioning result comprises: based on the matching result and the energy proportion of the target sound source signal in at least one sub-band, determining the target positioning result.
5. The method of claim 4, wherein, At least one sub-band comprises a first frequency band, and the first frequency band comprises a frequency band corresponding to human voice; The determination of the target positioning result based on the matching result and the energy proportion of the target sound source signal in at least one sub-band comprises: based on the matching result and the energy proportion of the target sound source signal in the first frequency band, determining a preliminary positioning result of the target sound source signal; based on the preliminary positioning result, determining the target positioning result.
6. The method of claim 5, wherein, The microphone array is arranged in the vehicle, and the reference phase difference is a phase difference of a reference sound source signal from inside the vehicle reaching different microphones in the microphone array, and different reference sound source signals are sound source signals from different preset positions inside the vehicle; The determination of the preliminary positioning result of the target sound source signal based on the matching result and the energy proportion of the target sound source signal in the first frequency band comprises: in response to the matching result representing that there is the matching phase difference, detecting whether the energy proportion of the first frequency band satisfies an energy threshold inside the vehicle; if the energy proportion of the first frequency band satisfies the energy threshold inside the vehicle, determining that the preliminary positioning result represents that the target sound source signal is derived from inside the vehicle and is from a preset position of the reference sound source signal corresponding to the matching phase difference; otherwise, determining that the preliminary positioning result represents that the target sound source signal is derived from outside the vehicle; in response to the matching result representing that there is no matching phase difference, determining that the preliminary positioning result represents that the target sound source signal is derived from outside the vehicle.
7. The method according to claim 5 or 6, characterized in that, The determination of the target positioning result based on the preliminary positioning result comprises: In response to the preliminary positioning result indicating that the target sound source signal originates from inside the vehicle, the target positioning result is determined to be that the target sound source signal originates from inside the vehicle; and / or, In response to the preliminary positioning result indicating that the target sound source signal originates from outside the vehicle, the preliminary positioning result is verified using the energy proportion of the target sound source signal in the second frequency band to obtain a corresponding verification result; the target positioning result is determined using the verification result, wherein the second frequency band is one of the sub-frequency bands in the at least one sub-frequency band, and the frequency of the second frequency band is higher than that of the first frequency band.
8. The method of claim 7, wherein, The preliminary positioning result is verified by utilizing the energy proportion of the target sound source signal in the second frequency band, resulting in a corresponding verification result, including: If the energy percentage of the second frequency band is not greater than a preset energy threshold, the verification result is determined to be correct for the preliminary positioning result; otherwise, the verification result is determined to be incorrect for the preliminary positioning result. The step of determining the target location result using the verification result includes: In response to the verification result indicating that the preliminary positioning result is correct, the target positioning result is determined to be that the target sound source signal originates from outside the vehicle. In response to the verification result indicating that the preliminary positioning result is incorrect, the target positioning result is determined to be that the target sound source signal originates from inside the vehicle.
9. The method of claim 1, wherein, After determining the target location result of the target sound source signal, the method further includes: in response to the target location result indicating that the target sound source signal originates from inside the vehicle, sending the target sound source signal to the vehicle's voice interaction module; in response to the target location result indicating that the target sound source signal originates from outside the vehicle, not responding to the target sound source signal; Acquiring the target sound source signal includes: acquiring the original sound source signal; determining the wake-up time based on the wake-up word or command word detected in the original sound source signal; and extracting the signal in the original sound source signal that is within a preset time range before and after the wake-up time as the target sound source signal.
10. A sound source positioning apparatus characterized by comprising: The device includes: The acquisition module is used to acquire the phase difference of the target sound source signal arriving at different microphones in the microphone array as the target phase difference, and to acquire the energy distribution data of the target sound source signal. The determination module is used to perform sound source localization by combining the target phase difference and the energy distribution data, and to determine the target localization result of the target sound source signal; the target localization result is used to characterize whether the target sound source signal originates from inside or outside the vehicle.
11. An electronic device, comprising: Including interconnected memory and processor, The memory stores program instructions; The processor is used to execute program instructions stored in the memory to implement the method according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that can be executed by a processor to implement the method of any one of claims 1-9.