A method, apparatus, device and storage medium for sound source localization

CN116027272BActive Publication Date: 2026-08-14AISPEECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]但是,一般的声源定位并不区分是否是人声,或者定位到声源后在对该声源进行语音增强后在判断是否是人声,当人声和干扰同时出现的时候,无法直接定位到人声,进而影响语音交互的效果

Benefits of technology

[0036] This invention provides a sound source localization method, apparatus, device, and storage medium. It performs speech separation on the original audio data, calculates the speech presence probability of each channel after separation, determines the speech channel containing the human voice based on the speech presence probability, and then applies the speech presence probability of the human voice channel back to the original audio data of multiple channels. This process extracts the human voice without affecting the phase relationship between channels. Finally, a sound source localization algorithm accurately locates the direction of the human voice. This achieves direct, accurate, and rapid location of the human voice during the sound source localization process, laying an accurate data foundation for subsequent voice interaction and thus improving the effectiveness of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116027272B_ABST
    Figure CN116027272B_ABST
Patent Text Reader

Abstract

This invention provides a sound source localization method, apparatus, device, and storage medium. The method includes: performing speech separation on received raw audio data to obtain multiple separated audio data streams; calculating the speech presence probability of each separated audio data stream; determining target separated audio data streams based on the speech presence probability of each stream; multiplying the speech presence probability corresponding to the target audio data streams by the raw audio data streams to obtain audio data to be localized; and performing sound source localization on the audio data to be localized to determine the location of human voices in the raw audio data streams. This invention enables direct, accurate, and rapid localization of human voices during the sound source localization process, laying an accurate data foundation for subsequent voice interaction and thus improving the effectiveness of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a sound source localization method, apparatus, device, and storage medium. Background Technology

[0002] With the advancement of technology, voice has become an increasingly popular means of human-computer interaction. However, voice interaction requires a microphone to collect voice information, but the voice collected by the microphone is always mixed with different random noises. This requires sound source localization of the sound collected by the microphone in order to identify accurate audio data.

[0003] However, typical sound source localization does not distinguish between human voices or not, or it performs speech enhancement on the sound source after localization before determining whether it is a human voice. When human voices and interference occur at the same time, it is impossible to directly localize the human voice, thus affecting the effect of voice interaction.

[0004] Therefore, how to provide a sound source localization method that can accurately locate human voices and improve the voice interaction effect is a technical problem that needs to be solved in this field. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a sound source localization method, apparatus, device, and storage medium to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of the present invention provides a sound source localization method, the method comprising the following steps:

[0007] The received raw audio data is subjected to speech separation to obtain multi-channel separated audio data;

[0008] Calculate the probability of speech presence in each of the multi-channel separated audio data;

[0009] The target separated audio data is determined based on the probability of speech presence in each of the separated audio data streams.

[0010] The probability of speech presence corresponding to the target audio separation data is multiplied by the original audio data to obtain the audio data to be located;

[0011] The source of the audio data to be located is determined by sound source localization, thus identifying the location of the human voice in the original audio data.

[0012] In some embodiments of the present invention, calculating the probability of speech presence in each of the multi-channel separated audio data includes:

[0013] The probability of speech presence in each of the multi-channel separated audio data is calculated using a neural network algorithm.

[0014] In some embodiments of the present invention, determining the target separated audio data based on the speech presence probability of each separated audio data stream includes:

[0015] Multiply the probability of speech presence in each separated audio data stream by the corresponding separated audio data to obtain the separated audio data to be processed;

[0016] Calculate the speech energy of the audio data to be processed and separate it, and select the audio data to be processed with the highest speech energy as the target audio data to be separated.

[0017] In some embodiments of the present invention, the step of performing sound source localization on the audio data to be located, and determining the location of human voices in the original audio data, includes:

[0018] Determine whether the probability of the speech corresponding to the target audio separation data is greater than a preset threshold. If it is greater, perform sound source localization on the audio data to be located to determine the location of the human voice in the original audio data.

[0019] If the probability of the speech corresponding to the target audio separation data is less than the preset threshold, then no sound source localization is performed.

[0020] Another aspect of the present invention provides a sound source localization device, the device comprising:

[0021] The audio separation module is used to perform speech separation on the received raw audio data to obtain multi-channel separated audio data;

[0022] The speech presence probability calculation module is used to calculate the speech presence probability of each of the multi-channel separated audio data.

[0023] The target audio determination module is used to determine the target separated audio data based on the speech presence probability of each separated audio data stream.

[0024] The audio processing module is used to multiply the probability of speech presence corresponding to the target audio separation data with the original audio data to obtain the audio data to be located;

[0025] The sound source localization module is used to locate the sound source of the audio data to be located and determine the location of the human voice in the original audio data.

[0026] In some embodiments of the present invention, the speech presence probability calculation module is specifically used for:

[0027] The probability of speech presence in each of the multi-channel separated audio data is calculated using a neural network algorithm.

[0028] In some embodiments of the present invention, the sound source localization module is specifically used for:

[0029] Determine whether the probability of the speech corresponding to the target audio separation data is greater than a preset threshold. If it is greater, perform sound source localization on the audio data to be located to determine the location of the human voice in the original audio data.

[0030] If the probability of the speech corresponding to the target audio separation data is less than the preset threshold, then no sound source localization is performed.

[0031] In some embodiments of the present invention, the target audio determination module is specifically used for:

[0032] Multiply the probability of speech presence in each separated audio data stream by the corresponding separated audio data to obtain the separated audio data to be processed;

[0033] Calculate the speech energy of the audio data to be processed and separate it, and select the audio data to be processed with the highest speech energy as the target audio data to be separated.

[0034] Another aspect of the present invention provides a sound source localization device, including a processor and a memory, wherein the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the above-described sound source localization method.

[0035] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described sound source localization method.

[0036] This invention provides a sound source localization method, apparatus, device, and storage medium. It performs speech separation on the original audio data, calculates the speech presence probability of each channel after separation, determines the speech channel containing the human voice based on the speech presence probability, and then applies the speech presence probability of the human voice channel back to the original audio data of multiple channels. This process extracts the human voice without affecting the phase relationship between channels. Finally, a sound source localization algorithm accurately locates the direction of the human voice. This achieves direct, accurate, and rapid location of the human voice during the sound source localization process, laying an accurate data foundation for subsequent voice interaction and thus improving the effectiveness of voice interaction.

[0037] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.

[0038] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description

[0039] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, are not intended to limit the scope of the invention. The components in the drawings are not drawn to scale but are merely illustrative of the principles of the invention. For ease of illustration and description of certain parts of the invention, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to the invention. In the drawings:

[0040] Figure 1 This is a schematic flowchart of a sound source localization method provided in one embodiment of this specification;

[0041] Figure 2 This is a schematic diagram of the sound source localization principle in one embodiment of this specification;

[0042] Figure 3 This is a schematic diagram of the module structure of one embodiment of the sound source localization device provided in this specification;

[0043] Figure 4 This is a hardware structure block diagram of a sound source localization server in one embodiment of this specification. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.

[0045] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.

[0046] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0047] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0048] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0049] With the advancement of technology, voice interaction has become widespread in people's work and daily lives. For example, users can communicate with friends and family far away through voice or video calls, and can also communicate with colleagues remotely through online meetings. Voice interaction requires the collection of audio data, followed by sound source localization to pinpoint the location of the sound and accurately capture the user's voice. However, in many scenarios, the user's environment during voice interaction is quite complex, resulting in a lot of noise in the collected audio data. Therefore, accurately locating the human voice is a technical problem that urgently needs to be solved.

[0050] Typical microphone sound source localization algorithms do not inherently distinguish between sound source types. They primarily estimate the location of one or more sound sources based on the time difference (phase difference) between the sound waves arriving at different microphones in the microphone array, thus failing to accurately pinpoint the location of a human voice. Therefore, to locate a specific sound source type, additional algorithms are needed to form a combined approach. However, even with a combination of algorithms, the general technique involves first performing localization, then enhancing the localized area with speech to determine if it is a human voice, rather than directly identifying and locating the human voice. When multiple noise sources are present, localization may be performed multiple times.

[0051] Existing sound source localization technologies focus on improving localization accuracy, and the current algorithms generally meet the accuracy requirements for most scenarios. However, when locating human voices in the presence of noise interference, multiple sound sources, including the noise source, are often located, or only one sound source is located, making it impossible to extract and locate the human voice separately. For scenarios with multiple noise sources, such as audio and video conferencing, it is necessary to exclude noise sources and accurately locate the human voice.

[0052] This specification provides a sound source localization method, mainly considering how to extract human voice in scenarios where human voice and noise coexist. This specification provides a method by separating the original audio data, identifying the channel where the human voice is located among the separated multi-channel audio, reflecting the human voice back onto the original audio, and then using a sound source localization algorithm to locate the sound source. This can accurately locate the location of the human voice, improve the accuracy of sound source localization, and thus provide an accurate data foundation for voice interaction.

[0053] Figure 1 This is a schematic flowchart of a sound source localization method provided in one embodiment of this specification, such as... Figure 1As shown in one embodiment of the sound source localization method provided in this specification, the method can be applied to terminal devices such as computers, tablets, servers, smartphones, and smart wearable devices. The method may include the following steps:

[0054] Step 102: Perform speech separation on the received raw audio data to obtain multi-channel separated audio data.

[0055] In practical implementation, during voice interaction, the voice interaction device receives raw audio data. Generally, this audio data can be received through a microphone array. A microphone array refers to an arrangement of microphones, that is, a system composed of a certain number of acoustic sensors (usually microphones) used to sample and process the spatial characteristics of the sound field. Typically, the raw audio data can be multiple audio signals received by the microphone array. In this embodiment, a speech separation algorithm is used to separate the received raw audio data, obtaining multiple separated audio data. At this point, the human voice is separated into one channel of the multiple separated audio data, and each channel has lost its original phase information, making location positioning impossible. Speech separation refers to the technique of separating mixed speech. The speech separation method can be selected according to actual needs, such as deep learning algorithms or deep clustering algorithms, to perform speech separation on the raw audio data. This embodiment does not specifically limit the method.

[0056] Step 104: Calculate the probability of speech presence in each of the multi-channel separated audio data.

[0057] In the specific implementation process, after obtaining multi-channel separated audio data by speech separation of the original audio data, the probability of speech presence in each channel of the multi-channel separated audio data can be calculated. The probability of speech presence can be understood as the probability that speech exists in a channel. Generally, some channels in the multi-channel separated audio data may not contain speech, and the channels containing human voices are generally the channels with the highest probability of speech presence. Intelligent learning algorithms can be used to calculate the probability of speech presence in the multi-channel separated audio data, or other methods can be used for calculation. This specification does not specifically limit the embodiments.

[0058] In some embodiments of this specification, calculating the probability of speech presence in each of the multi-channel separated audio data includes:

[0059] The probability of speech presence in each of the multi-channel separated audio data is calculated using a neural network algorithm.

[0060] In the specific implementation process, the embodiments of this specification can use neural network algorithms to estimate the probability of the existence of single-channel audio speech. When there is interference, the network estimation performance will decrease. However, by using neural network algorithms to calculate the probability of the existence of speech in each channel after speech separation, the probability of the existence of human voice speech can be estimated more accurately in the separated human voice channel, thus laying an accurate data foundation for subsequent human voice localization.

[0061] Step 106: Determine the target separated audio data based on the speech presence probability of each separated audio data stream.

[0062] In the specific implementation process, the target separated audio data can be understood as the audio channel where the human voice is located. That is, after calculating the probability of speech existence in each channel of the multi-channel separated audio data, the channel where the human voice is located is determined according to the probability of speech existence in each channel of the separated audio data, and the target separated audio data is obtained.

[0063] In some embodiments of this specification, determining the target separated audio data based on the speech presence probability of each separated audio data stream includes:

[0064] Multiply the probability of speech presence in each separated audio data stream by the corresponding separated audio data to obtain the separated audio data to be processed;

[0065] Calculate the speech energy of the audio data to be processed and separate it, and select the audio data to be processed with the highest speech energy as the target audio data to be separated.

[0066] In the specific implementation process, after calculating the speech presence probability of the multi-channel separated audio data, the speech presence probability of each channel of separated audio data can be multiplied by the corresponding separated audio data to obtain the separated audio data to be processed. For example, after speech separation of the original audio data, three channels of separated audio data Y1, Y2, and Y3 are obtained. The speech presence probabilities corresponding to the three channels of separated audio data Y1, Y2, and Y3 are p1, p2, and p3, respectively. Multiplying each channel of separated audio data by its corresponding speech presence probability yields: Y1×p1, Y2×p2, and Y3×p3, resulting in three channels of separated audio data to be processed. Then, the speech energy of each channel of separated audio data to be processed is calculated, and the channel of separated audio data with the highest speech energy is taken as the target separated audio data, i.e., the audio data of the channel containing the human voice. The method for calculating speech energy can be selected according to actual needs. For example, the square of the average speech amplitude in each channel of separated audio data to be processed can be calculated, or other methods can be used for calculation. This specification does not specifically limit the method used in the embodiments.

[0067] By multiplying the probability of speech presence with the corresponding audio data and then calculating the speech energy of the processed audio data, the channel where the human voice is located can be quickly and accurately located, laying an accurate data foundation for subsequent human voice localization.

[0068] Step 108: Multiply the probability of speech presence corresponding to the target audio separation data by the original audio data to obtain the audio data to be located.

[0069] In the specific implementation process, after determining the speech channel where the human voice is located, the probability of speech presence in that channel is used as a feedback effect on the original audio data. For example, multiplying the probability of speech presence corresponding to the target audio separation data with the original audio data yields the phase-invariant multi-channel audio of the human voice, i.e., the audio data to be located. This extracts the human voice without affecting the phase relationship between channels.

[0070] Step 110: Perform sound source localization on the audio data to be located to determine the location of human voices in the original audio data.

[0071] In the specific implementation process, the probability of the presence of speech in the channel containing the human voice is applied back to the original audio data. Then, a sound source localization algorithm is used to locate the sound source in the processed audio data to be located. This directly locates the location of the human voice in the original audio data, and suppresses noise, indirectly improving the accuracy of sound source localization. The location of the human voice can be understood as the location of the speaker or main speaker in the original audio data. Generally, voice interaction scenarios may contain a lot of noise, such as ambient noise or other human voice interference. The embodiments in this specification mainly aim to locate the location of the main speaker in the voice interaction scenario. For example, in scenarios such as audio conferencing, multiple people may be conducting an audio conference on one voice interaction device, but generally only one speaker speaks at a time. However, when the speaker speaks, other users may also make noise. This requires accurately locating the speaker's location to identify accurate audio data. Of course, in some scenarios, the audio data collected by voice interaction may include other noise, such as the sound of a television or a washing machine. The main purpose of the embodiments in this specification is to locate the location of the person speaking in the voice interaction.

[0072] The method for sound source localization can be selected according to actual needs. For example, it can be a controllable beamforming technique based on maximum output power, a high-resolution spectral estimation technique, or a sound source localization technique based on sound time difference. The embodiments in this specification do not make specific limitations.

[0073] The sound source localization method provided in this specification's embodiments separates the original audio data into speech segments, calculates the speech presence probability of each channel after separation, determines the speech channel containing the human voice based on the speech presence probability, and then applies the speech presence probability of the human voice channel back to the original audio data of multiple channels. This process extracts the human voice without affecting the phase relationship between channels. Finally, a sound source localization algorithm accurately locates the direction of the human voice. This method achieves direct, accurate, and rapid location of the human voice during the sound source localization process, laying an accurate data foundation for subsequent voice interaction and thus improving the effectiveness of voice interaction.

[0074] In some embodiments of this specification, the step of locating the sound source of the audio data to be located and determining the location of the human voice in the original audio data includes:

[0075] Determine whether the probability of the speech corresponding to the target audio separation data is greater than a preset threshold. If it is greater, perform sound source localization on the audio data to be located to determine the location of the human voice in the original audio data.

[0076] If the probability of the speech corresponding to the target audio separation data is less than the preset threshold, then no sound source localization is performed.

[0077] In the specific implementation process, when locating the sound source of the audio data to be located, it can be first determined whether the probability of the presence of speech in the channel where the human voice is located, i.e., the probability of the presence of speech corresponding to the target audio separation data, is greater than a preset threshold. If it is greater, it indicates that there is human voice in the audio of the current voice interaction, and a sound source localization algorithm can be used to locate the sound source of the audio data to be located, determining the location of the human voice in the original audio data. If the probability of the presence of speech corresponding to the target audio separation data is less than the preset threshold, it indicates that the currently collected original audio data does not include human voice, and sound source localization can be skipped, directly ignoring the collected audio data, or treating the collected original audio data as noise. The size of the preset threshold can be set according to actual needs, and this embodiment does not impose specific limitations.

[0078] In the localization process described in this specification, a threshold is set based on the probability of voice presence to determine whether to locate the current audio. This ensures that sound source localization is only performed when human voice is detected, avoiding unnecessary data processing, ensuring system performance, and improving the data processing efficiency of sound source localization.

[0079] Furthermore, in some embodiments of this specification, after receiving the original audio data, noise reduction processing can be performed on the original audio data, and then a neural network algorithm can be used to calculate the probability of speech presence in the noise-reduced audio data. The noise reduction method can be selected according to actual needs, and this specification does not impose specific limitations on the embodiments. That is, depending on the actual usage, a noise reduction algorithm can be used to replace the speech separation technology in the embodiments of this specification, distinguishing noise from human voice. Noise reduction processing of the audio data before calculating the probability of speech presence can improve the accuracy of speech presence probability calculation, thereby improving the accuracy of human voice localization in the sound source.

[0080] Figure 2 This is a schematic diagram of the sound source localization principle in one embodiment of this specification, as shown below. Figure 2 As shown, the main process of sound source localization in the embodiments of this specification can be referred to as follows:

[0081] Step 1: The microphone array receives the original multi-channel signals and uses a speech separation algorithm to separate the original signals to obtain multiple audio channels.

[0082] Step 2: Each output path of the speech separation algorithm is processed by a neural network to obtain the speech presence probability of each path. The neural network outputs the speech presence probability (0-1) at each frequency point in the current frame in the frequency domain, as shown in the following formula:

[0083] P = (p1, p2, ..., p) F ), 0≤p i ≤1

[0084] The probability of speech presence in each separated audio stream can be the sum of the probabilities of speech presence at all frequency points in that audio stream.

[0085] Step 3: Select and output the speech presence probability of the channel containing the human voice. Multiply the speech presence probability output by the neural network by the corresponding audio paths after speech separation to obtain the speech of each path. Calculate the speech energy of each path, select the path with the highest energy as the speech of the channel containing the human voice, and output the speech presence probability of that path.

[0086] Step 4: Multiply the probability of the output speech by the original audio received by the microphone array to obtain phase-invariant multi-channel audio of human voice.

[0087] Step 5: Set a threshold δ. When the probability of speech in the selected human voice channel is greater than this threshold, sound source localization is performed, thereby achieving the localization of human voices independently in noisy scenarios. If the probability is less than this threshold, sound source localization is not performed.

[0088] This embodiment uses a speech separation algorithm to separate the original audio, resulting in multi-channel audio, including human voice. However, the human voice is now separated into a single channel, and each channel has lost its original phase information, making location positioning impossible. Therefore, it is necessary to consider using the separated human voice information to feed back into the original audio. This embodiment uses a neural network algorithm to estimate the probability of speech presence in each separated channel. Compared to directly using a neural network algorithm to estimate the original audio, this yields a more accurate probability of human voice presence. After speech separation and processing with the neural network outputting the probability of speech presence, the human voice is extracted without affecting the phase relationship between channels. Then, a sound source localization algorithm is used to locate the human voice, suppressing noise and indirectly improving localization accuracy. Furthermore, a speech presence probability threshold is added, ensuring that localization is only output when human voice is present. This embodiment combines multiple algorithms to extract human voice in noisy environments without affecting phase, and combined with a localization algorithm for human voice localization, achieves accurate human voice localization in voice interaction scenarios.

[0089] The various embodiments of the methods described in this specification are presented in a progressive manner. Similar or identical parts between the embodiments can be referred to interchangeably. Each embodiment focuses on highlighting the differences from other embodiments. Relevant details can be found in the descriptions of the method embodiments.

[0090] Based on the sound source localization method described above, one or more embodiments of this specification also provide a sound source localization apparatus. The apparatus may include devices (including distributed systems), software (applications), modules, components, servers, clients, etc., using the methods described in the embodiments of this specification, combined with necessary hardware implementations. Based on the same innovative concept, the apparatuses in one or more embodiments provided in this specification are as described in the following embodiments. Since the implementation schemes and methods for solving the problem are similar, the implementation of specific apparatuses in the embodiments of this specification can refer to the implementation of the foregoing methods, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the apparatuses described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.

[0091] Specifically, Figure 3 This is a schematic diagram of the module structure of one embodiment of the sound source localization device provided in this specification, as shown below. Figure 3 As shown, the apparatus provided in this specification may include:

[0092] The audio separation module 31 is used to perform speech separation on the received raw audio data to obtain multi-channel separated audio data;

[0093] The speech presence probability calculation module 32 is used to calculate the speech presence probability of each of the multi-channel separated audio data.

[0094] The target audio determination module 33 is used to determine the target separated audio data based on the speech presence probability of each separated audio data stream;

[0095] The audio processing module 34 is used to multiply the probability of speech presence corresponding to the target audio separation data with the original audio data to obtain the audio data to be located;

[0096] The sound source localization module 35 is used to locate the sound source of the audio data to be located and determine the location of the human voice in the original audio data.

[0097] In some embodiments of this specification, the speech presence probability calculation module is specifically used for:

[0098] The probability of speech presence in each of the multi-channel separated audio data is calculated using a neural network algorithm.

[0099] In some embodiments of this specification, the sound source localization module is specifically used for:

[0100] Determine whether the probability of the speech corresponding to the target audio separation data is greater than a preset threshold. If it is greater, perform sound source localization on the audio data to be located to determine the location of the human voice in the original audio data.

[0101] If the probability of the speech corresponding to the target audio separation data is less than the preset threshold, then no sound source localization is performed.

[0102] In some embodiments of this specification, the target audio determination module is specifically used for:

[0103] Multiply the probability of speech presence in each separated audio data stream by the corresponding separated audio data to obtain the separated audio data to be processed;

[0104] Calculate the speech energy of the audio data to be processed and separate it, and select the audio data to be processed with the highest speech energy as the target audio data to be separated.

[0105] The sound source localization device provided in the embodiments of this specification can extract human voices in noisy scenes without affecting the phase. Combined with the use of a localization algorithm for human voice localization, it achieves accurate localization of human voices in voice interaction scenarios.

[0106] In some embodiments of this specification, a sound source localization device is also provided, including a processor and a memory. The memory stores computer instructions, and the processor executes the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the sound source localization method described in the above embodiments, such as:

[0107] The received raw audio data is subjected to speech separation to obtain multi-channel separated audio data;

[0108] Calculate the probability of speech presence in each of the multi-channel separated audio data;

[0109] The target separated audio data is determined based on the probability of speech presence in each of the separated audio data streams.

[0110] The probability of speech presence corresponding to the target audio separation data is multiplied by the original audio data to obtain the audio data to be located;

[0111] The source of the audio data to be located is determined by sound source localization, thus identifying the location of the human voice in the original audio data.

[0112] It should be noted that the apparatus and equipment described above, based on the method embodiments, may also include other implementation methods. Specific implementation methods can be found in the descriptions of the relevant method embodiments, and will not be elaborated upon here.

[0113] The methods and embodiments provided in this specification can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Taking execution on a server as an example... Figure 4 This is a hardware structure block diagram of a sound source localization server in one embodiment of this specification. The computer terminal can be the sound source localization server or sound source localization device in the above embodiment. Figure 4 The server 10 shown may include one or more (only one is shown in the figure) processors 100 (processors 100 may include, but are not limited to, microprocessors MCUs or programmable logic devices FPGAs), non-volatile memory 200 for storing data, and a transmission module 300 for communication functions. Those skilled in the art will understand that... Figure 4 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 10 may also include components that are more... Figure 4 The components shown may include more or fewer components, and may also include other processing hardware such as databases or multi-level caches, GPUs, or components with similar capabilities. Figure 4 The different configurations shown.

[0114] The non-volatile memory 200 can be used to store software programs and modules of application software, such as the program instructions / modules corresponding to the sound source localization method in the embodiments of this specification. The processor 100 executes various functional applications and resource data updates by running the software programs and modules stored in the non-volatile memory 200. The non-volatile memory 200 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the non-volatile memory 200 may further include memory remotely located relative to the processor 100, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0115] The transmission module 300 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computer terminal's communication provider. In one example, the transmission module 300 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 300 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0116] Corresponding to the above method, the present invention also provides an apparatus comprising a computer device including a processor and a memory, the memory storing computer instructions, the processor executing the computer instructions stored in the memory, and the apparatus performing the steps of the method as described above when the computer instructions are executed by the processor.

[0117] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0118] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.

[0119] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.

[0120] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0121] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for locating a sound source, characterized in that, The method includes: The received raw audio data is subjected to speech separation to obtain multi-channel separated audio data; The probability of speech presence at each frequency point in the current frame of each of the multi-channel separated audio data is calculated using a neural network algorithm. Multiply the probability of speech presence at each frequency point in the current frame of each separated audio data by the corresponding separated audio data to obtain the separated audio data to be processed; Calculate the speech energy of the audio data to be separated, and take the audio data to be separated with the highest speech energy as the target audio separation data. The probability of speech presence corresponding to the target audio separation data is multiplied by the original audio data to obtain the audio data to be located; The source of the audio data to be located is determined by sound source localization, thus identifying the location of the human voice in the original audio data.

2. The method according to claim 1, characterized in that, The step of locating the sound source of the audio data to be located, and determining the location of the human voice in the original audio data, includes: Determine whether the probability of the speech corresponding to the target audio separation data is greater than a preset threshold. If it is greater, perform sound source localization on the audio data to be located to determine the location of the human voice in the original audio data. If the probability of the speech corresponding to the target audio separation data is less than the preset threshold, then no sound source localization is performed.

3. A sound source localization device, characterized in that, The device includes: The audio separation module is used to perform speech separation on the received raw audio data to obtain multi-channel separated audio data; The speech presence probability calculation module is used to calculate the speech presence probability at each frequency point in the current frame of each of the multi-channel separated audio data using a neural network algorithm. The target audio determination module is used to multiply the probability of speech presence at each frequency point in the current frame of each separated audio data by the corresponding separated audio data to obtain the separated audio data to be processed; calculate the speech energy of the separated audio data to be processed, and take the separated audio data to be processed with the largest speech energy as the target audio separation data; The audio processing module is used to multiply the probability of speech presence corresponding to the target audio separation data with the original audio data to obtain the audio data to be located; The sound source localization module is used to locate the sound source of the audio data to be located and determine the location of the human voice in the original audio data.

4. The apparatus according to claim 3, characterized in that, The sound source localization module is specifically used for: Determine whether the probability of the speech corresponding to the target audio separation data is greater than a preset threshold. If it is greater, perform sound source localization on the audio data to be located to determine the location of the human voice in the original audio data. If the probability of the speech corresponding to the target audio separation data is less than the preset threshold, then no sound source localization is performed.

5. A sound source localization device, comprising a processor and a memory, characterized in that, The memory stores computer instructions, and the processor executes the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method as described in any one of claims 1 to 2.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Speech recognition method, intelligent equipment and intelligent television

    CN110875045A

  • Voice signal processing method, device and equipment and storage medium

    CN111883166A

  • Voice data processing method and device, electronic equipment and storage medium

    CN113870841A