Sound source positioning model training method, sound source object positioning method, and related apparatus

By using a sound source localization model trained through multiple rounds of iterations, the unit content vector and audio content vector of multi-channel audio signals are extracted, and the position prediction vector is adjusted. This solves the problems of inaccurate sound source object localization and excessive resource consumption in existing technologies, and achieves efficient and accurate sound source object localization.

WO2026026320A1PCT designated stage Publication Date: 2026-02-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102956
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2025-06-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing technologies suffer from incomplete speech separation during sound source object localization, leading to loss of speech information and inaccurate recognition of wake words. Furthermore, multi-wake word recognition models consume excessive resources, reducing resource utilization and localization efficiency.

Method used

A sound source localization model trained through multiple rounds of iterations is adopted. By extracting the unit content vector and audio content vector of multi-channel audio signals, adjusting the position prediction vector, and combining the correlation degree, the predicted position of the sound source object is determined, which simplifies the localization process and improves localization accuracy and resource utilization efficiency.

Benefits of technology

This technology enables accurate location of sound source objects in multi-channel audio signals, improving location accuracy, reducing resource consumption, simplifying processing logic, and enhancing the efficiency and accuracy of sound source object location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102956_05022026_PF_FP_ABST
    Figure CN2025102956_05022026_PF_FP_ABST
Patent Text Reader

Abstract

A sound source positioning model training method, comprising using training samples to perform a plurality of rounds of iterative training on a constructed initial sound source positioning model, wherein one round of iterative process comprises: a control device reads a training sample (301), the training sample comprising a multi-channel sample audio signal marked with a sample wake-up word, and a location tag corresponding to the multi-channel sample audio signal; the control device respectively extracts, for pronunciation units comprised in the sample wake-up word, unit content vectors respectively corresponding to the pronunciation units, and extracts, according to the multi-channel sample audio signal, a location prediction vector fused with audio content information, and an audio content vector fused with location information of an acquisition subject (302), the location prediction vector being used for describing spatial locations respectively corresponding to at least one acquisition subject having a voice expression; the control device adjusts the location prediction vector respectively on the basis of the degrees of association between the unit content vectors and the audio content vector, to obtain sound source location indication vectors of the pronunciation units respectively corresponding to the unit content vectors, and determines a predicted location of a sound source object of the sample wake-up word on the basis of the obtained sound source location indication vectors (303); and the control device adjusts model parameters of a sound source positioning model on the basis of the result difference between the predicted position and a corresponding location tag (304).
Need to check novelty before this filing date? Find Prior Art

Description

Training methods for sound source localization models, methods for locating sound source objects, and related devices.

[0001] Related applications

[0002] This application claims priority to Chinese patent application filed on July 31, 2024, with application number 202411034743.0, entitled "Training method for sound source localization model, sound source object localization method and related device", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of data processing technology, and in particular to a training method for a sound source localization model, a sound source object localization method, and related apparatus. Background Technology

[0004] In the current technology, in order to accurately respond to the voice requests of objects in different locations, it is necessary to locate the sound source object of the voice-triggered candidate wake word.

[0005] For example, in a vehicle environment, in order to meet the operational needs of different people in different seats, the vehicle device can locate and respond to the sound source object after the sound source object expresses the wake word.

[0006] Currently, when locating a sound source, the process typically involves first performing speech separation and noise reduction on the multi-channel audio signal according to preset candidate positions. Then, a wake-up word recognition model is used for each candidate position to identify audio signals belonging to different candidate positions. The candidate position corresponding to the audio signal containing the candidate wake-up word is then determined as the location of the sound source.

[0007] However, this localization method is highly susceptible to the effects of speech separation. If the speech content is not completely separated, the audio signal of a candidate position will include speech interference from other positions. Moreover, important speech information may be lost in the audio signal after noise reduction and suppression, making it impossible to effectively recognize wake words, thus making it impossible to locate the sound source object and accurately respond to the needs of the sound source object. In addition, using multiple wake word recognition models requires a lot of processing resources, thereby reducing resource utilization. Summary of the Invention

[0008] This application provides a method for training a sound source localization model, a method for locating sound source objects, and related devices.

[0009] Firstly, a training method for a sound source localization model is proposed, executed by an electronic device. Using various training samples, the initial sound source localization model is trained through multiple iterations. Each iteration includes:

[0010] Read training samples; the training samples include: multi-channel sample audio signals labeled with sample wake words, and position labels corresponding to the multi-channel sample audio signals;

[0011] For each articulation unit contained in the sample wake word, the unit content vector corresponding to each articulation unit is extracted. Based on the multi-channel sample audio signal, a position prediction vector fused with audio content information and an audio content vector fused with the position information of each collected object are extracted. The position prediction vector is used to describe the spatial position of at least one collected object with speech expression.

[0012] Based on the degree of correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted respectively to obtain the sound source position indication vector of the pronunciation unit corresponding to each unit content vector, and based on the obtained sound source position indication vector, the predicted position of the sound source object of the sample wake word is determined.

[0013] Based on the difference between the predicted location and the corresponding location label, the model parameters of the sound source localization model are adjusted.

[0014] Secondly, a method for locating a sound source object is proposed, executed by a control device, including:

[0015] Acquire multi-channel audio signals from the audio acquisition component;

[0016] When identifying a target wake word that successfully matches each of the preset candidate wake words based on the multi-channel audio signal, the trained target sound source localization model is used to extract the unit content vector corresponding to each pronunciation unit contained in the target wake word. Based on the multi-channel audio signal, a position prediction vector fused with audio content information and an audio content vector fused with the position information of the collected object are extracted. The position prediction vector is used to describe the spatial position of at least one collected object with speech expression.

[0017] Based on the correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of the pronunciation unit corresponding to each unit content vector, and the location of the sound source object of the target wake word is determined based on the obtained sound source position indication vector.

[0018] Thirdly, a training device for a sound source localization model is proposed. The device includes a training unit for performing multiple rounds of iterative training on the constructed initial sound source localization model using various training samples. One round of iterative training includes:

[0019] Read training samples; the training samples include: multi-channel sample audio signals labeled with sample wake words, and position labels corresponding to the multi-channel sample audio signals;

[0020] For each articulation unit contained in the sample wake word, the unit content vector corresponding to each articulation unit is extracted. Based on the multi-channel sample audio signal, a position prediction vector fused with audio content information and an audio content vector fused with the position information of each collected object are extracted. The position prediction vector is used to describe the spatial position of at least one collected object with speech expression.

[0021] Based on the degree of correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted respectively to obtain the sound source position indication vector of the pronunciation unit corresponding to each unit content vector, and based on the obtained sound source position indication vector, the predicted position of the sound source object of the sample wake word is determined.

[0022] Based on the difference between the predicted location and the corresponding location label, the model parameters of the sound source localization model are adjusted.

[0023] Fourthly, a sound source object localization device is proposed, comprising:

[0024] The acquisition unit is used to acquire multi-channel audio signals collected by the audio acquisition component;

[0025] The localization unit, when identifying a target wake-up word that successfully matches each of the preset candidate wake-up words based on the multi-channel audio signal, uses a trained target sound source localization model to extract the unit content vector corresponding to each pronunciation unit contained in the target wake-up word, and extracts a position prediction vector fused with audio content information and an audio content vector fused with the position information of the collected object based on the multi-channel audio signal. The position prediction vector is used to describe the spatial position corresponding to at least one collected object with speech expression. Based on the correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of the pronunciation unit corresponding to each unit content vector, and the localization position of the sound source object of the target wake-up word is determined based on the obtained sound source position indication vectors.

[0026] Fifthly, an electronic device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.

[0027] In a sixth aspect, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the above-described method.

[0028] In a seventh aspect, a computer program product is proposed, comprising a computer program that, when executed by a processor, implements the above-described method.

[0029] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the published drawings without creative effort.

[0031] Figure 1 is a schematic diagram of the sound source object positioning process under the relevant technology in the embodiments of this application;

[0032] Figure 2 is a schematic diagram of possible application scenarios in the embodiments of this application;

[0033] Figure 3A is a schematic diagram of the model structure of the sound source localization model in the embodiment of this application;

[0034] Figure 3B is a schematic diagram of the spatial model constructed by simulation in the embodiment of this application;

[0035] Figure 3C is a schematic diagram of a training process of the sound source localization model in an embodiment of this application;

[0036] Figure 4A is a schematic diagram of the sound source object positioning process in an embodiment of this application;

[0037] Figure 4B is a schematic diagram of the model processing process in an embodiment of this application;

[0038] Figure 5 is a schematic diagram of the vector change process during the location of the sound source object in the embodiments of this application;

[0039] Figure 6 is a schematic diagram of the feature dimension changes inside the target sound source localization model in the embodiment of this application;

[0040] Figure 7 is a schematic diagram of the logical structure of the training device for the sound source localization model in the embodiment of this application;

[0041] Figure 8 is a schematic diagram of the logical structure of the sound source object positioning device in an embodiment of this application;

[0042] Figure 9 is a schematic diagram of the hardware composition structure of an electronic device applying an embodiment of this application;

[0043] Figure 10 is a schematic diagram of the hardware composition structure of another electronic device using an embodiment of this application. Detailed Implementation

[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0045] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in sequences other than those illustrated or described herein.

[0046] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0047] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.

[0048] Direction of arrival (DOA): This usually refers to the angle of a sound source (or sound source) relative to an audio acquisition device (e.g., a microphone) used to acquire its audio signal. In this embodiment, the sound source localization can also be understood as predicting the DOA of the sound source. Specifically, the DOA can be predicted by training an initial sound source localization model. At this time, for each candidate position determined for the sound source, specifically, it refers to the various possible spatial angles of the sound source relative to its corresponding microphone. Moreover, the predicted spatial angles can reflect the spatial position of the corresponding sound source in the localization environment.

[0049] Orientation confidence: also known as prediction probability, is used to evaluate the position matching score of the sound source object at each candidate position, which is determined by mapping each position indicator vector. Specifically, taking the processing based on a position indicator vector as an example, the model can internally map the position indicator vector determined for a corresponding speech unit to obtain the prediction probability at each candidate position.

[0050] Sound source object: In this embodiment of the application, the object that expresses the wake word by speech is called the sound source object; it should be noted that, in combination with the actual processing situation, this application assumes that there is at most one sound source object of the candidate wake word at the same time.

[0051] The object being acquired refers to the object whose audio signal is being acquired, that is, the object with speech expression that is acquired by the audio acquisition component in the space where the audio acquisition component is deployed; among them, the sound source object can be regarded as an object being acquired.

[0052] Candidate wake-up words: In this embodiment of the application, it refers to preset words that can evoke a targeted response from the control device.

[0053] Joint front-end and back-end modeling refers to using a common objective function to simultaneously optimize the model parameters of the speech front-end model and the speech back-end model. Here, the speech front-end refers to the feature vector extraction process, and the speech back-end refers to the localization and prediction process.

[0054] Impulse response: refers to the response generated in response to a unit impulse signal.

[0055] The design concept of the embodiments of this application is briefly introduced below:

[0056] In the current technology, in order to accurately respond to the voice requests of objects in different locations, it is necessary to locate the sound source object of the voice-triggered candidate wake word.

[0057] For example, in a vehicle environment, in order to meet the operational needs of different people in different seats, the vehicle device can locate and respond to the sound source object after the sound source object expresses the wake-up word, and handle the processing needs of the sound source object in a targeted manner.

[0058] Currently, when locating a sound source, the process typically involves first acquiring audio signals from various candidate locations using multiple microphones deployed in a fixed location. Then, an acoustic front-end separates and denoises the acquired multi-channel audio signals according to preset candidate locations, obtaining the audio signals corresponding to each candidate location. Subsequently, a wake-up word recognition model, preset for each candidate location, is used to determine whether the corresponding audio signal contains a preset candidate wake-up word. The candidate location corresponding to the wake-up word recognition model that detects the candidate wake-up word is taken as the location of the sound source.

[0059] Referring to Figure 1, which is a schematic diagram of the localization process of the sound source object under the relevant technology in the embodiment of this application, after the speech separation front-end (i.e., the acoustic front-end) acquires the multi-channel audio signal (or multi-channel speech signal), assuming that there are N preset candidate directions (or candidate positions), the audio signal of each direction is separated from the multi-channel audio signal. Then, the wake word recognition model corresponding to each candidate direction is used to simultaneously perform wake word recognition on the audio signals of the N candidate directions to obtain the wake word recognition result corresponding to each candidate direction. Further, using post-processing logic, based on the N wake word recognition results, the candidate direction corresponding to the wake word recognition model that identifies the candidate wake word is determined as the localization direction of the sound source object.

[0060] However, this processing method is highly dependent on the performance of the acoustic front-end. If the acoustic front-end fails to completely separate the audio signals from different candidate locations, the audio signal at one candidate location will include audio interference from other candidate locations. This results in multiple location positions being determined for a single sound source object after processing by the wake-word recognition model. Consequently, in the post-processing logic, further processing combining audio energy and acoustic scores is required to obtain the final location position, which significantly reduces the localization efficiency of the sound source object. Moreover, the need to start multiple wake-word recognition models increases the consumption of processing resources. To improve response speed, the performance of the wake word recognition model usually needs to be downgraded, which greatly affects the model's recognition accuracy. In addition, since the optimization goal of the acoustic front end is to include only the speech of the acquired object in the "auditory perception", the high-frequency and low-frequency content in the audio signals at each candidate position will be processed, which will suppress certain frequency bands. As a result, the processed audio signal cannot meet the recognition requirements of the wake word recognition model, so it cannot effectively recognize the wake word, thus it cannot locate the sound source object, and it cannot accurately respond to the needs of the sound source object, and thus cannot guarantee the recognition effect of the model.

[0061] In view of this, this application proposes a training method, a sound source localization method, and related apparatus for a sound source localization model. During the training of the sound source localization model, multiple rounds of iterative training are performed on the constructed initial sound source localization model using various training samples. Each iteration includes: reading training samples, which include: multi-channel sample audio signals labeled with sample wake words, and position labels corresponding to the multi-channel sample audio signals; then, for each articulation unit contained in the sample wake word, extracting the unit content vector corresponding to each articulation unit, and extracting the fused audio content vector based on the multi-channel sample audio signals. The system comprises a location prediction vector for information and an audio content vector incorporating the location information of the collected object. The location prediction vector describes the spatial location of at least one collected object with a speech expression. Then, based on the correlation between each unit content vector and the audio content vector, the location prediction vector is adjusted to obtain the sound source location indication vector of the pronunciation unit corresponding to each unit content vector. Based on the obtained sound source location indication vectors, the predicted location of the sound source object of the sample wake word is determined. Finally, based on the difference between the predicted location and the corresponding location label, the model parameters of the sound source localization model are adjusted.

[0062] Thus, during the training of the initial sound source localization model, by establishing the correlation between each unit content vector and the audio content vector, the positions of vectors with high correlation to the corresponding unit content vectors can be determined within the audio content vectors. Then, based on the correlation, the position prediction vectors incorporating audio content information are adjusted, enabling the determination of sound source position indication vectors representing the position of the corresponding sound source object from the perspective of each articulatory unit. Furthermore, based on each sound source position indication vector, the predicted position of the sound source object can be mapped and determined. This transforms the sound source object localization problem into addressing the localization of each speech unit representing the sound source object. The system maps the positions of the lines, and simultaneously trains the vector extraction and position prediction processes throughout the overall training process. The training objective is to improve the accuracy of the predicted positions, achieving end-to-end modeling of the localization process, which helps to obtain better model training results. In addition, the system can directly process multi-channel sample audio data during the overall training process without needing to perform position separation and noise reduction on the multi-channel sample audio data. This simplifies the processing logic of the localization process, enabling the system to learn the ability to locate sound source objects at any candidate position based on multi-channel sample speech signals and included sample wake words, and significantly improves the accuracy of localization.

[0063] In the process of sound source object localization, multi-channel audio signals acquired by the audio acquisition component are first obtained. Then, when the target wake-up word that successfully matches each of the preset candidate wake-up words is identified based on the multi-channel audio signals, the trained target sound source localization model is used to perform sound source object localization processing. This makes the trigger condition for the sound source localization process to be the detection of the speech expression of the candidate wake-up word in the multi-channel audio signal. Moreover, when using the target sound source localization model for specific processing, the following operations are performed: for each articulation unit contained in the target wake-up word, the unit content vector corresponding to each articulation unit is extracted, and based on the multi-channel audio signal, the position prediction vector fused with audio content information and the audio content vector fused with the position information of the acquired object are extracted. The position prediction vector is used to describe the spatial position of each acquired object with speech expression. Then, based on the degree of correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of the articulation unit corresponding to each unit content vector, and based on the obtained sound source position indication vectors, the localization position of the sound source object of the target wake-up word is determined.

[0064] In this way, on the one hand, the target sound source localization model can be used to locate the sound source object at the appropriate processing time, enabling accurate localization of the sound source object expressing the target wake word under the wake word. On the other hand, when determining the corresponding sound source position indication vector for each articulatory unit, the overall extracted position prediction vector can be adaptively adjusted based on the correlation between the unit content vector corresponding to the articulatory unit and the audio content vector that incorporates the position information of the collected object. This allows for the generation of a sound source position indication vector describing the position of the corresponding sound source object for each articulatory unit. Therefore, the position of the sound source object can be predicted from the perspective of each articulatory unit expressing the sound source object, thereby improving the accuracy of the sound source object localization process. Moreover, thanks to the excellent processing performance of the target sound source localization model, the amount of resources consumed in the sound source object localization process can be reduced, improving resource utilization efficiency.

[0065] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0066] Referring to Figure 2, which is a schematic diagram of a possible application scenario in an embodiment of this application, the schematic diagram includes a controlled device 210 and a control device 220.

[0067] In some feasible embodiments of this application, the control device 220 can receive multi-channel audio signals acquired in real time by the audio acquisition component within the target space. The audio acquisition component may include multiple microphones; each microphone can acquire the audio signal of one channel; the target space includes preset candidate positions. The total number of microphones included in the audio acquisition component may be the same as or different from the total number of candidate positions; this application does not impose specific limitations on this. Furthermore, during the real-time wake-up word recognition process of the received multi-channel audio signals, when the control device 220 determines that the multi-channel audio signals include a target wake-up word that successfully matches one of the preset candidate wake-up words, it begins to locate the sound source object for the multi-channel audio signals containing the target wake-up word.

[0068] In some other feasible embodiments of this application, after the control device 220 completes the positioning of the sound source object, it can respond specifically to the sound source object at the positioning location, and adjust the controlled device 210 associated with the sound source object according to further instructions from the sound source object.

[0069] Controlled devices 210 include, but are not limited to, mobile phones, tablets, laptops, e-book readers, smart voice interaction devices, smart home appliances, in-vehicle functional devices (such as displays, switches that control window opening and closing), aircraft, etc.

[0070] Control devices 220 include, but are not limited to, mobile phones, tablets, laptops, e-book readers, smart voice interaction devices, central control devices deployed in smart home systems, vehicle terminals, aircraft, etc.; or, they can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0071] In this embodiment, the controlled device 210 and the control device 220 can communicate via a wired network or a wireless network. The following description focuses solely on the control device 220, illustrating the training process of the sound source localization model and the sound source object localization process.

[0072] The following is an illustrative description of possible sound source object location scenarios:

[0073] Application Scenario 1: In an in-vehicle environment, determine the seat location of the sound source object that triggers the wake word.

[0074] The control device can preset candidate positions (e.g., driver's seat, front passenger seat, left rear passenger, and right rear passenger) for the in-vehicle environment. Multiple microphones are deployed in the vehicle in a compact or separate manner. Then, according to the actual processing needs, the control device can collect audio signals from the passengers in each candidate position. Based on the actual processing needs, when the collected multi-channel audio signals include the target wake-up word, the device can locate the sound source object that triggered the target wake-up word based on the multi-channel audio signals, and determine the location of the sound source object.

[0075] Furthermore, a targeted response is provided to the sound source object at the location, and a targeted response to the audio signal at the location is activated. In addition, the controlled device associated with the location of the sound source object is adjusted according to the processing instructions of the sound source object.

[0076] For example, after the control device detects that the front passenger has triggered the target wake-up word, it can control the speaker to respond with "Hello, front passenger", and then respond to the front passenger's further instructions.

[0077] Application Scenario 2: In a smart home scenario, determine the location of the home environment where the sound source object that triggers the wake word is located.

[0078] The control device can preset candidate locations for each smart home environment (e.g., living room area, master bedroom area, secondary bedroom area, kitchen area); then, according to actual processing needs, the control device can collect audio signals from the users in each candidate location, and, based on actual processing needs, determine when the collected multi-channel audio signals include the target wake-up word, locate the sound source object that triggers the target wake-up word based on the multi-channel audio signals, and determine the location of the sound source object.

[0079] Furthermore, it provides a targeted response to the sound source object at the location, activates a targeted response to the audio signal at the location, and adjusts the smart home devices associated with the location of the sound source object according to the processing instructions of the sound source object.

[0080] Application Scenario 3: In security and maintenance scenarios, determine the location of the sound source object that triggers the wake word.

[0081] The control equipment can preset candidate locations (e.g., location A, area B, area C, etc.) for the security and maintenance environment. Then, according to the actual processing needs, the control equipment can collect audio signals from the security and maintenance personnel in each candidate location. Based on the actual processing needs, if the collected multi-channel audio signals include the target wake-up word, it can be determined that the security and maintenance personnel in a certain venue area have triggered the target wake-up word. Then, based on the multi-channel audio signals, the sound source object that triggered the target wake-up word is located, and the location of the sound source object is determined.

[0082] Furthermore, it provides targeted responses to the sound source at the location, activates targeted responses to the audio signals at the location, and adjusts controllable equipment within the venue area where the sound source is located according to the sound source's instructions.

[0083] In addition, it should be understood that the specific implementation of this application involves the location of the sound source object and the training of the sound source location model. When the embodiments described in this application are applied to specific products or technologies, the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0084] The training process of the sound source localization model will be explained below from the perspective of the control equipment, with reference to the accompanying diagram:

[0085] Referring to Figure 3A, which is a schematic diagram of the model structure of the sound source localization model in an embodiment of this application, the model structure of the constructed initial sound source localization model will be described below with reference to Figure 3A:

[0086] As shown in Figure 3A, the initial sound source localization model includes: a two-dimensional multi-channel convolutional network, a speech recognition encoder network, a speech feature extraction encoder network, a speech unit encoding network, a self-attention network, an attention network, and fully connected layers.

[0087] Two-dimensional multi-channel convolutional network: capable of extracting beam features; during training, it is used to extract comprehensive multi-channel information based on multi-channel sample audio signals to obtain a mixed audio vector.

[0088] The speech recognition encoder network is used to extract features from speech content based on mixed audio vectors, obtaining a content description vector. The network structure can be a Factorized Sparse Matrix Network (FSMN); a Long Short-Term Memory (LSTM) network; or a sequence model network combining the advantages of Transformer Networks and Convolutional Neural Networks (CNNs). The speech recognition encoder network in the initial sound source localization model can be obtained after pre-training, depending on the actual processing needs. The Factorized Sparse Matrix Network is a network structure that can be used in speech recognition encoder networks to extract features from speech content based on mixed audio vectors, obtaining a content description vector. The Transformer Network is a neural network architecture based on an attention mechanism. In this application, it can be combined with Convolutional Neural Networks (CNNs) to form a sequence model network, serving as an optional network structure for the speech recognition encoder network to extract the content description vector from the mixed audio vectors.

[0089] The speech feature extraction encoder network is used to extract a location description vector representing the spatial location of the acquired object based on a mixture of audio vectors. The network structure can specifically be a two-dimensional convolutional network. The speech recognition encoder network in the initial sound source localization model can be obtained after initialization, depending on the actual processing needs. The location description vector refers to the vector representing the spatial location of the acquired object extracted by the speech feature extraction encoder network based on the mixture of audio vectors.

[0090] Speech Unit Encoding Network: During training, this network extracts the corresponding unit content vector for each articulation unit included in the sample wake word. The network structure can be: encoding network + self-attention network; or, only the encoding network can be used for processing. This application does not impose any specific restrictions on this.

[0091] Attention network: refers to a network built with the help of attention mechanisms.

[0092] Self-attention networks are networks built using self-attention mechanisms, such as those built using the self-attention mechanism in Transformers.

[0093] The process of training the sound source localization model will be explained below with reference to the accompanying diagram:

[0094] In this embodiment, the example of the control device training an initial sound source localization model to obtain a target sound source localization model is used for illustrative purposes. During the training process of the initial sound source localization model, the control device first constructs training samples, and then performs multiple rounds of iterative training on the constructed initial sound source localization model based on the training samples until a preset convergence condition is met, thereby obtaining the trained target sound source localization model. The preset convergence condition may be: the number of training rounds reaches a first set value, or the number of times the model loss value calculated after a single round of training is continuously lower than a second set value reaches a third set value. The values ​​of the first set value, the second set value, and the third set value are set according to the actual processing needs, and this application does not impose specific restrictions on them.

[0095] The construction process of each training sample will be explained below:

[0096] It should be noted that, in this embodiment of the application, considering the influence of the varying number of audio acquisition devices, multi-channel audio signals are difficult to obtain from actual applications, making it impossible to flexibly construct corresponding training samples for each application scenario. Therefore, this application can use an offline simulation method to construct multi-channel sample audio signals and construct training samples based on the simulated multi-channel sample audio signals for model training. During the simulation process, the emission position of the sound source signal (i.e., impulse signal) needs to be recorded as the position label in the corresponding generated training samples. The specific offline simulation method is as follows: First, based on the specific business scenario when locating the sound source object, computer simulation software (such as COMSOL Multiphysics) is used to simulate and construct the corresponding spatial model based on the geometric shape, acoustic characteristics, and other parameters of the scenario. For simulated candidate positions, discrete position points can be divided at certain intervals according to the area in the scenario where sound source objects may exist as candidate positions. For simulated audio acquisition positions, the corresponding positions can be determined in the spatial model according to the deployment method and range of the actual audio acquisition devices. Then, for each candidate position, the following operations are performed: a unit impulse signal is emitted at the candidate position, and the impulse response results at each audio acquisition position are calculated using the finite element method. The specific steps are as follows: First, an acoustic model is established based on parameters such as the geometry and acoustic characteristics of the scene, and the model is divided into a finite number of elements. Then, the stiffness matrix and mass matrix of each element are constructed. Next, the matrix equation of the entire system is assembled according to the acoustic wave equation and boundary conditions. Finally, the matrix equation is solved to obtain the impulse response results at each audio acquisition location. The impulse response results at each audio acquisition location are arranged in channel order to obtain the multi-channel impulse response.

[0097] In this embodiment of the application, each training sample includes: a multi-channel sample audio signal labeled with the sample wake word, and a position label corresponding to the multi-channel sample audio signal.

[0098] Multi-channel sample audio signal refers to the audio signal simulated by combining the target sound source localization model with the specific application scenario when constructing training samples, determining the total number of channels based on the total number of audio acquisition devices capable of acquiring audio signals, and simulating the audio signal by convolutional fusion of single-channel audio signals with the target impulse response. It can be a time domain signal or a frequency domain signal, and is used to train the initial sound source localization model.

[0099] Location labels are labels selected from candidate locations of the sound source object in a sound source localization scenario, corresponding to the multi-channel sample audio signal. They are used to mark the spatial location of the sound source object and are compared with the sound source object location predicted by the model during the training of the sound source localization model in order to adjust the model parameters.

[0100] When constructing multi-channel sample audio signals, the control device needs to determine the total number of audio acquisition devices (such as the total number of microphones) capable of acquiring audio signals, based on the specific application scenario of the trained target sound source localization model, and then determine the total number of multi-channels based on the total number of audio acquisition devices. Additionally, it should be noted that since audio signal analysis typically involves the frequency domain, the time-domain signal needs to be converted to the frequency-domain signal before processing with the model. Therefore, it should be clear that one microphone can acquire the time-domain audio signal of one channel. After processing the time-domain audio signal of one channel through a Fast Fourier Transform, two frequency-domain signals can be obtained. Considering that frequency-domain signals are usually in complex form, the two frequency-domain signals represent the real and imaginary parts obtained from the transformation of a single time-domain signal.

[0101] Based on this, depending on the actual processing needs, the multi-channel sample audio signal included in each training sample can be a multi-channel time-domain audio signal, or the multi-channel sample audio signal included in each training sample can be a multi-channel frequency-domain audio signal. This application does not impose any specific restrictions on this.

[0102] Taking the construction of a multi-channel sample audio signal in a training sample as an example, the control device can select a target impulse response from the multi-channel impulse responses associated with the label positions respectively. Among them, a multi-channel impulse response includes: the response result of an impulse signal sent from a label position, which is determined by simulation for multiple preset audio acquisition positions in the sound source localization scenario; a label position is selected from the candidate positions of the sound source object in the sound source localization scenario; then a single-channel audio signal expressing a sample wake word is acquired, and the corresponding multi-channel sample audio signal is simulated based on the single-channel audio signal and the target impulse response.

[0103] Multi-channel impulse response refers to the response results obtained by simulating and determining the response of multiple preset audio acquisition positions in a sound source localization scenario to an impulse signal sent from a tag position, arranged in channel order. It is used to simulate the reception of a single-channel audio signal sent from a candidate position by each audio acquisition position, so as to construct a multi-channel sample audio signal.

[0104] It should be noted that, in this embodiment, in order to simulate the impact of the audio acquisition component on the audio signals at different candidate locations, and to flexibly construct training samples according to the actual sample generation needs, the control device needs to pre-construct the multi-channel impulse responses associated with the label locations. Specifically, when constructing the multi-channel impulse responses, the control device can first use computer simulation software (such as COMSOL Multiphysics) to simulate and construct the corresponding spatial model based on the specific business scenario when locating the sound source object, according to the geometric shape, acoustic characteristics, and other parameters of the scenario. For simulated candidate locations, discrete location points can be divided into candidate locations at certain intervals according to the areas in the scenario where sound source objects may exist; for simulated audio acquisition locations, the corresponding locations can be determined in the spatial model according to the deployment method and range of the actual audio acquisition equipment. Then, for each candidate location, the following operations are performed: a unit impulse signal is emitted at the candidate location, and the impulse response results at each audio acquisition location are calculated using the finite element method. The specific steps are as follows: First, an acoustic model is established based on parameters such as the scene's geometry and acoustic characteristics, and the model is divided into a finite number of elements. Then, the stiffness matrix and mass matrix of each element are constructed. Next, the matrix equation of the entire system is assembled based on the acoustic wave equation and boundary conditions. Finally, the matrix equation is solved to obtain the impulse response results at each audio acquisition location. The impulse response results at each audio acquisition location are arranged in channel order to obtain the multi-channel impulse response, and then the multi-channel impulse response with the candidate location as the label location is obtained.

[0105] For example, referring to Figure 3B, which is a schematic diagram of the simulated spatial model constructed in an embodiment of this application, it is assumed that the actual sound source object positioning scenario is: locating the sound source object in a car carrying four people. Based on this, a cuboid can be constructed to simulate the vehicle space, and then the four seat positions can be used as candidate positions. When selecting a candidate position, the candidate position can be selected as slightly above the seat. At the same time, according to the actual deployment method and position of the microphones in the vehicle, each audio acquisition position can be marked in the vehicle space. The deployment method of each microphone can be a compact deployment or a separate deployment, and this application does not impose specific restrictions on this.

[0106] Furthermore, based on the actual training sample generation needs, the control device selects the target impulse response from among the multi-channel impulse responses to construct a training sample. In the process of simulating the corresponding multi-channel sample audio signal based on a single-channel audio signal and the target impulse response, the control device can convolve and fuse the target impulse response with a single-channel audio signal to obtain a multi-channel simulated audio signal. Then, this multi-channel simulated audio signal is determined as the multi-channel sample audio signal.

[0107] A multi-channel analog audio signal is a time-domain signal obtained by convolving and fusing the target impulse response with a single-channel audio signal. It can be determined as a multi-channel sample audio signal according to actual processing needs. The multi-channel sample audio signal can be a time-domain signal or a frequency-domain signal.

[0108] It should be noted that when the multi-channel analog audio signal is a time-domain signal, after convolutional fusion processing, a multi-channel time-domain audio signal can be directly obtained. Then, according to actual processing needs, the multi-channel time-domain audio signal can be processed to obtain a multi-channel sample audio signal. This multi-channel sample audio signal can be a time-domain signal or a frequency-domain signal; this application does not impose specific limitations on this. Specifically, when the constructed multi-channel sample audio signal is a time-domain signal, the multi-channel analog audio signal in time-domain form can be directly used as the multi-channel sample audio signal. Similarly, when the constructed multi-channel sample audio signal is a frequency-domain signal, it is necessary to process the multi-channel analog audio signal to ultimately obtain a multi-channel sample audio signal that meets the requirements.

[0109] In addition, the selected single-channel audio signal contains the speech expression of a wake-up word. Then, based on the single-channel audio signal and combined with the selected target impulse response, the corresponding multi-channel sample audio signal is simulated. The wake-up word can be used as the sample wake-up word determined based on the multi-channel sample audio signal, and the position label corresponding to the selected target impulse response can be used as the position label associated with the multi-channel sample audio signal. Then, based on the multi-channel sample audio signal and its associated position label and sample wake-up word, a corresponding training sample is constructed.

[0110] In this way, by convolving and fusing the single-channel audio signal with the selected target impulse response, a multi-channel sample audio signal is obtained. With the help of the multi-channel sample audio signal, it is possible to simulate the reception of the single-channel audio signal emitted from a candidate position at each audio acquisition position, thereby realizing the flexible construction of the multi-channel sample audio signal used for training and improving the construction efficiency of training samples.

[0111] Similarly, the processing device can construct multiple training samples based on each single-channel audio signal containing a wake-up word and each multi-channel impulse response. The wake-up words included in different single-channel audio signals may differ; the total number of multi-channel impulse responses is the same as the total number of candidate positions preset according to the actual application scenario. Containing a wake-up word means that the audio signal contains speech content expressing the wake-up word.

[0112] In this way, by constructing corresponding multi-channel impulse responses for each candidate position, the signal reception at each audio acquisition position can be concretely described by using the simulated response at each audio acquisition position. Furthermore, under the action of the selected target impulse response, the reception of single-channel audio signals at each audio acquisition position can be simulated, thereby flexibly determining multi-channel sample audio signals and constructing suitable training samples according to any actual application scenario, which greatly reduces the difficulty of constructing training samples.

[0113] Optionally, in this embodiment, considering that in actual business scenarios, there may be more than one audio source expressing the wake-up word simultaneously, and other objects being collected in the same space may also be expressing their voices. Moreover, the collected audio signal is inevitably affected by ambient noise. Therefore, in order to make the constructed multi-channel sample audio signal more closely resemble the actual collection scenario, it is necessary to perform noise addition processing on the multi-channel sample audio signal.

[0114] In the specific processing, the control device can use any one or a combination of the following noise-adding methods, including but not limited to: after simulating and obtaining the multi-channel sample audio signal in a training sample, the multi-channel sample audio signal is subjected to noise-adding processing:

[0115] Noise addition method 1: Superimpose environmental noise.

[0116] Specifically, the control device directly acquires the collected multi-channel environmental noise and superimposes the acquired multi-channel environmental noise onto the multi-channel sample audio signal, wherein the multi-channel sample audio signal and the multi-channel environmental noise have the same total number of channels.

[0117] It should be noted that multi-channel ambient noise refers to noise signals directly acquired using audio acquisition components in a specific application scenario. These noise signals are superimposed on the multi-channel sample audio signals during training sample construction to simulate the impact of ambient noise on the audio signals in the actual acquisition scenario. The total number of channels in multi-channel ambient noise is the same as that in the multi-channel sample audio signals. Multi-channel ambient noise can be directly acquired using audio acquisition components in a specific application scenario.

[0118] For example, in a vehicle-mounted scenario, multi-channel ambient noise is specifically collected by multiple audio acquisition components within the vehicle.

[0119] Noise addition method 2: Superimpose interference signals with known differences.

[0120] Specifically, the control device can construct a multi-channel interference signal with the same total number of channels as the multi-channel sample audio signal, and superimpose the multi-channel interference signal onto the multi-channel sample audio signal.

[0121] In this embodiment of the application, when constructing the corresponding multi-channel interference signal for the multi-channel sample audio signal, the following two feasible construction methods can be adopted (but are not limited to these two):

[0122] Method 1: The multi-channel audio signal simulated for objects with speech expression in other locations is identified as the multi-channel interference signal.

[0123] Multi-channel interference signal refers to a signal with the same total number of channels as the multi-channel sample audio signal, which is constructed in a specific way (such as determining the multi-channel audio signal simulated for objects with speech expression at other locations as interference signals, or obtaining the interference signal by configuring the signal-to-noise ratio and combining it with a preset base signal according to the distribution of multiple audio acquisition locations) when constructing training samples in order to simulate interference in the actual audio acquisition process. It is used to superimpose the multi-channel sample audio signal.

[0124] Specifically, the control device can select an interference location among the various tag locations, acquire the multi-channel impulse response at the interference location, and simulate the corresponding multi-channel interference signal based on the multi-channel impulse response at the interference location and the single-channel audio signal that does not express the sample wake-up word.

[0125] It should be noted that in the embodiments of this application, the selected interference position can be one or more. When there are multiple interference positions, the corresponding multi-channel interference signal can be simulated based on the multi-channel impulse response corresponding to each interference position and combined with the preset single-channel audio signal that does not express the sample wake-up word. When generating multi-channel interference signals corresponding to different interference positions, the single-channel audio signal used can be the same or different. This application does not impose specific restrictions on this.

[0126] Method 2: Based on the distribution of multiple audio acquisition locations, configure the corresponding signal-to-noise ratio for each audio acquisition location, and obtain a multi-channel interference signal based on the obtained signal-to-noise ratio and the preset base signal.

[0127] In this embodiment of the application, in order to adapt to the differences in signal-to-noise ratio of audio signals collected from different audio collection locations due to the distance between some audio collection locations and the sound source object in actual scenarios, the corresponding signal-to-noise ratio can be configured with a certain probability for different channels in the multi-channel sample audio signal according to the actual processing needs. The probability value is set according to the actual processing needs.

[0128] Specifically, the control device can configure the corresponding signal-to-noise ratio for each channel in the multi-channel sample audio signal according to the distribution of multiple audio acquisition locations, and obtain the corresponding multi-channel interference signal based on the signal-to-noise ratio of each channel and the preset base signal.

[0129] Specifically, the base signal is a multi-channel noise signal, and the base signal and the multi-channel sample audio signal have the same total number of channels. When obtaining the multi-channel interference signal, the corresponding signal-to-noise ratio can be used to process the signal content of each channel in the base signal, and finally the multi-channel interference signal can be obtained.

[0130] In this way, by using the multi-channel interference signal construction methods indicated in Construction Method 1 and Construction Method 2, we can better simulate the interference in the actual audio acquisition process, making the constructed training samples more in line with the signal state in the actual scenario and improving the construction effect of the training samples.

[0131] Noise addition method three: perform phase delay processing.

[0132] In this embodiment of the application, considering that in actual business processes, due to the quality differences between different audio acquisition devices, the multi-channel audio signals acquired at different audio acquisition locations usually have phase consistency differences; therefore, the control device can perturb the phase of the multi-channel sample audio signals with a certain probability to simulate the delay of the signal arriving at the audio acquisition device (such as a microphone) and simulate the interference caused by the quality differences between audio acquisition devices.

[0133] Specifically, the control device can perform phase delay processing on multi-channel sample audio signals according to the preset phase delay levels for each of the corresponding channels. The phase delay levels for each of the multiple channels can be randomly set according to actual processing needs. Furthermore, the phase delay can be performed in the time domain or the frequency domain. This application does not impose specific restrictions on the method used for phase delay processing.

[0134] In this way, by using noise addition methods one through three, perturbations can be added to the obtained multi-channel sample audio signals to take into account the interference that may exist in the actual scene. Training samples are then constructed based on the perturbated multi-channel sample audio signals, so that the obtained training samples can simulate the interference in the real scene. Furthermore, when training the model based on the perturbated training samples, the robustness of the model can be enhanced, which helps the model learn the ability to locate sound source objects in high noise and strong interference environments.

[0135] Furthermore, referring to Figure 3C, which is a schematic diagram of one round of training of the sound source localization model in an embodiment of this application, the following describes the operations performed in one round of training of the initial sound source localization model, taking one round of iterative training as an example:

[0136] Step 301: Control the device to read the training samples.

[0137] The training samples include: multi-channel sample audio signals labeled with sample wake words, and position labels corresponding to the multi-channel sample audio signals.

[0138] Step 302: The control device extracts the unit content vector corresponding to each pronunciation unit contained in the sample wake word, and extracts the position prediction vector fused with audio content information and the audio content vector fused with the position information of the sampled object based on the multi-channel sample audio signal.

[0139] The location prediction vector is used to describe the spatial location of each of the at least one subject with speech expression; the at least one subject with speech expression refers to the object whose audio signal is acquired by an audio acquisition device at at least one audio acquisition location; for each articulation unit included in the sample wake-up word, it can be any one of elements such as phonemes, syllables, or characters, and this application does not impose any specific restrictions on it.

[0140] Unit content vector refers to the vector obtained by analyzing and identifying the pronunciation units contained in the sample wake word or target wake word, encoding the initial vectors, and then using a self-attention mechanism to capture the correlation between the initial vectors. This vector is used to subsequently determine the sound source location indicator vector of the pronunciation unit.

[0141] Audio content vectors are vectors that are obtained by processing multi-channel sample audio signals or multi-channel audio signals through a series of processes (such as extracting content description vectors and location description vectors, and superimposing and mapping to obtain a comprehensive information vector). They are vectors that incorporate the location information of each collected object and are used as K vectors in attention networks.

[0142] In this embodiment, the control device uses an initial sound source localization model. When extracting the unit content vector corresponding to each pronunciation unit for each pronunciation unit contained in the sample wake-up word, in a feasible implementation, the pronunciation units contained in the sample wake-up word can be analyzed and determined first, and then feature extraction can be performed directly on each pronunciation unit to obtain the corresponding unit content vector. In other feasible implementations, the pronunciation units contained in the sample wake-up word can be analyzed and determined, and the corresponding initial vector can be encoded for each pronunciation unit. Then, a self-attention mechanism can be used to capture the correlation between the initial vectors to obtain the unit content vector determined for each pronunciation unit.

[0143] It should be noted that when using the self-attention mechanism for processing, (Query, Q) vector, (Key, K) vector, and (Value, V) vector can be predicted and generated based on the initial vector of each articulatory unit. Then, by calculating the similarity between the Q vector and the K vector, the correlation between each initial vector is determined, and a weight matrix is ​​obtained. Finally, the V vector is weighted and fused based on the obtained weight matrix to obtain the unit content vector determined for each articulatory unit.

[0144] Specifically, when processing using a self-attention mechanism, for each articulatory unit, the initial vector x... i (i = 1, 2, ..., n, where n is the number of articulatory units), through three different linear transformation matrices W Q W K and W V Perform linear transformations on each vector to predict and generate the (Query, Q) vector q. i =W Q x i (Key, K) vector k i =W K x i and the (Value, V) vector v i =W V x i Then, by calculating the dot product similarity between the Q vector and the K vector, i.e. And the similarity is scaled (usually divided by ). d k (where K is the dimension of the vector), and then normalized using the softmax function to obtain the weight matrix A, where... Then, the V vector is weighted and fused based on the obtained weight matrix, that is... The unit content vectors corresponding to each articulation unit are obtained.

[0145] In this way, by leveraging the self-attention mechanism, when extracting the corresponding unit content vector for each articulatory unit, the influence of articulatory units in other positions can be integrated, making the obtained articulatory units more suitable for actual use needs and improving the extraction effect of unit content vectors.

[0146] Meanwhile, in the process of extracting the location prediction vector fused with audio content information and the audio content vector fused with the location information of each collected object based on the multi-channel sample audio signal, the control device can extract the content description vector representing the speech content and the location description vector representing the spatial location of each collected object based on the multi-channel sample audio signal; then superimpose the content description vector and the location description vector mapped to the same vector space to obtain the comprehensive information vector; then, based on the comprehensive information vector, extract the location prediction vector fused with audio content information and the audio content vector fused with the location information of the collected object.

[0147] Specifically, the control device can use a two-dimensional multi-channel convolutional network to perform convolutional fusion processing on multi-channel sample audio signals. The two-dimensional multi-channel convolutional network contains multiple convolutional layers and pooling layers. In the convolutional layers, appropriate kernel sizes and strides are set to extract local features of the multi-channel audio signals through convolution operations. In the pooling layers, appropriate pooling types (such as max pooling or average pooling) are used, and appropriate pooling window sizes and strides are set to perform feature dimensionality reduction through pooling operations, ultimately obtaining a mixed audio vector. To extract the content description vector from the mixed audio vector, the mixed audio vector can be input into a speech recognition encoder network (such as a Factorized Sparse Matrix (FSMN)). In FSMN, linear transformations use matrix multiplication, i.e., the input vector is multiplied by the weight matrix and then the bias vector is added; non-linear activation functions use common activation functions (such as ReLU, Sigmoid, etc.). Through a series of such linear transformations and non-linear activation functions, the overall corresponding content description vector is extracted from the speech content level. To extract location description vectors from mixed audio vectors, the mixed audio vectors can be input into a speech feature extraction encoder network (such as a 2D convolutional network). In the 2D convolutional network, the convolutional layers are configured with appropriate kernel size and stride to extract features related to the spatial location of each captured object from the mixed audio vectors through convolution operations. The pooling layers employ appropriate pooling types (such as max pooling or average pooling), with appropriate pooling window size and stride, to perform feature dimensionality reduction, extracting the overall location description vector from the spatial location of each captured object. Furthermore, to fuse the content-level vectors and location-level features to obtain a comprehensive information vector, a fully connected layer can be used for dimensionality unification. By setting an appropriate number of neurons in the fully connected layer, the content description vectors and location description vectors are mapped to a vector space of the same dimension, and then the vectors are superimposed to achieve feature fusion. Next, two fully connected layers are set up, each with an appropriate number of neurons. The first fully connected layer maps the comprehensive information vector to an audio content vector that incorporates location information; the second fully connected layer maps the comprehensive information vector to a location prediction vector that incorporates audio content information.

[0148] Among them, the integrated information vector refers to the vector obtained by superimposing the content description vector representing the speech content and the location description vector representing the spatial location of each collected object after mapping them to the same vector space through a fully connected layer. It is used to extract the location prediction vector that incorporates audio content information and the audio content vector that incorporates the location information of the collected objects.

[0149] The location prediction vector can be understood as a latent vector extracted based on the phase and energy characteristics of a multi-channel audio signal, or as a vector representation of the phase and energy information of a multi-channel audio signal, which can characterize the spatial location information of each collected object.

[0150] Hybrid audio vectors refer to vectors obtained by a two-dimensional multi-channel convolutional network extracting comprehensive multi-channel information from multi-channel sample audio signals during the training of a sound source localization model. These vectors can then be used to further extract content description vectors and location description vectors.

[0151] This allows for audio signal-level processing, enabling the extraction of corresponding audio content vectors and location prediction vectors from the perspectives of audio content and the location of the acquired object, based on multi-channel sample speech signals. By leveraging vector superposition and extraction during feature extraction, the speaker's location information can be incorporated into the extracted audio content vector, and the audio content information can be incorporated into the extracted location prediction vector. This strengthens the connection between feature vectors from different perspectives and helps improve the construction effect when building location description vectors for speech units.

[0152] Step 303: The control device adjusts the position prediction vector based on the correlation between each unit content vector and the audio content vector, obtains the sound source position indication vector of the corresponding pronunciation unit for each unit content vector, and determines the predicted position of the sound source object of the sample wake word based on the obtained sound source position indication vector.

[0153] The sound source position indication vector is a vector that represents the position of the sound source object corresponding to the corresponding sound unit. It is obtained by adjusting the position prediction vector based on the correlation between each unit content vector and the audio content vector. The predicted position or localization position of the sound source object can be determined by mapping each sound source position indication vector.

[0154] Predicted position refers to the process of performing vector mapping on each obtained sound source position indication vector during the training of the sound source localization model to obtain the mapped position and the corresponding predicted probability. The mapped positions whose predicted probabilities meet the first preset condition are selected, and the mapped positions whose total occurrences meet the second preset condition are used as the positions determined for the sound source objects of the sample wake words. These positions are then compared with the position labels to adjust the model parameters.

[0155] Mapping location refers to the process of performing vector mapping on each obtained sound source location indicator vector during the training of the sound source localization model or the localization of the sound source object. After obtaining the predicted probability of each candidate location, the target location whose predicted probability meets the first preset condition is selected and used to further determine the predicted location or localization location of the sound source object.

[0156] Specifically, within the initial sound source localization model, the content vectors of each unit can be used as the Q vector in the attention network, the determined audio content vector can be used as the K vector in the attention network, and the position prediction vector can be used as the V vector in the attention network. This allows the position prediction vector to be adjusted based on the degree of correlation between each unit content vector and the audio content vector, thereby obtaining the sound source position indication vector of the corresponding pronunciation unit for each unit content vector.

[0157] Furthermore, when determining the predicted position of the sound source object of the sample wake word based on the obtained sound source position indication vectors, vector mapping processing can be performed on each sound source position indication vector to obtain the mapped position and the corresponding predicted probability. The mapped position is the target position whose predicted probability satisfies the first preset condition after obtaining the predicted probability under each candidate position. Then, the mapped position whose total occurrence satisfies the second preset condition is used as the predicted position determined for the sound source object of the sample wake word.

[0158] Specifically, when using the initial sound source localization model for processing, a fully connected layer can be used to perform vector mapping processing based on the obtained sound source location indication vectors to obtain the prediction probability at each candidate location. Then, for the prediction probability at each candidate location corresponding to each sound source location indication vector, the mapping location whose prediction probability satisfies the first preset condition can be determined. The first preset condition can be: the corresponding prediction probability is the highest. After obtaining each mapping location associated with the prediction probability, each mapping location can be filtered according to the second preset condition to obtain the prediction location determined for the sound source object of the sample wake word. The second preset condition can be: the total number of occurrences is the highest. The total number of occurrences is determined by statistically analyzing the occurrence of each mapping location.

[0159] When processing using the initial sound source localization model, for each sound source location indicator vector z i (i = 1, 2, ..., m, where m is the number of sound source location indicator vectors), and input them into a fully connected layer. The fully connected layer contains a weight matrix W and a bias vector b, which are linearly transformed by u... i =Wz i +b yields the intermediate vector u i Then for the intermediate vector u i Apply the softmax function, that is (j = 1, 2, ..., N, where N is the number of candidate positions), obtain the predicted probability p at each candidate position. ijFurthermore, for each candidate position corresponding to the sound source position indicator vector, the mapping position whose prediction probability satisfies the first preset condition is determined. The first preset condition may be: the corresponding prediction probability is the largest. After obtaining each mapping position associated with the prediction probability, each mapping position can be filtered according to the second preset condition to obtain the prediction position determined for the sound source object of the sample wake word. The second preset condition may be: the total number of occurrences is the largest. The total number of occurrences is determined by statistically analyzing the occurrence of each mapping position.

[0160] In this way, after determining the sound source location indicator vector from the perspective of each speech unit, the corresponding mapping position can be obtained based on the sound source location indicator vector. Then, by combining the mapping positions, the final predicted position can be determined. This is equivalent to breaking down the sound source object localization problem and performing position prediction separately from each speech unit expressed by the sound source object, which can improve the accuracy of sound source object localization.

[0161] Step 304: The control device adjusts the model parameters of the sound source localization model based on the result difference between the predicted location and the corresponding location label.

[0162] Specifically, after obtaining the predicted position, the control device can adjust the model parameters based on the difference between the predicted position and the position labels corresponding to the read training samples, using a focal loss function. The expression for the focal loss function is FL(p t )=-(1-p t ) γ log(p t ), where p t Here, γ represents the probability predicted by the model, and γ is an adjustment factor. During training, the difference between the predicted location and the location label is first calculated and substituted into the focus loss function to calculate the loss value. Then, the backpropagation algorithm is used to update the model parameters based on the loss value. For example, stochastic gradient descent (SGD) can be used, with the update formula being: Where, θ old These are the current model parameter values, θ new These are the updated model parameter values, and α is the learning rate, which controls the step size for each parameter update. This describes the gradient of the loss function FL with respect to the model parameters θ. The backpropagation algorithm, in neural network training, refers to the algorithm that, based on the loss value calculated using the loss function, propagates the error back from the output layer to the input layer to update the model parameters. In this application, it is used to update the parameters of the sound source localization model based on the loss value calculated using the focus loss function.

[0163] Similarly, the control device can perform multiple rounds of iterative training according to the single-round training process indicated in steps 301-304, and finally obtain the trained target sound source localization model.

[0164] The process of locating a sound source object based on the target sound source localization model is explained below with reference to the accompanying diagram:

[0165] Referring to Figure 4A, which is a schematic diagram of the sound source object positioning process in an embodiment of this application, the process of positioning the sound source object is described below with reference to Figure 4A:

[0166] Step 401: Control the device to acquire the multi-channel audio signals acquired by the audio acquisition component.

[0167] In actual processing, with the authorization of the relevant object, the control device can control the audio acquisition component to acquire multi-channel audio signals in real time, thereby enabling the control device to acquire the multi-channel audio signals acquired by the audio acquisition component in real time; or, after the relevant object actively enables the audio recognition function, the control device can acquire the multi-channel audio signals acquired by the audio acquisition component in real time while the audio recognition function is enabled.

[0168] The multi-channel audio acquisition component includes multiple audio acquisition devices deployed in a compact or separate manner. Optionally, depending on the actual processing needs, the real-time acquired audio data can be divided into audio segments with a specified number of frames, and then the sound source object can be located based on the multi-channel audio signal corresponding to each audio segment. The value of the specified number of frames is set according to the actual processing needs, and this application does not impose specific restrictions on it.

[0169] Step 402: When the control device identifies the target wake-up word that successfully matches each preset candidate wake-up word based on the multi-channel audio signal, it uses the trained target sound source localization model to identify the location of the sound source object.

[0170] In this context, "location" refers to the process of locating the sound source object by using a trained target sound source localization model to adjust the location prediction vector based on the correlation between the content vector of each unit and the audio content vector, thereby obtaining the sound source location indication vector of each pronunciation unit. The location of the sound source object of the target wake word is determined based on the sound source location indication vector, which is then used for subsequent targeted response and processing of the sound source object.

[0171] The target wake-up word refers to the word that is successfully matched with each of the preset candidate wake-up words in the multi-channel audio signal acquired by the audio acquisition component. When the target wake-up word is identified, the trained target sound source localization model is triggered to locate the sound source object.

[0172] In this embodiment, when the control device identifies the target wake word from multi-channel audio signals, one feasible implementation is to use a trained multi-channel wake word recognition model. The multi-channel audio signals are input into a Feedforward Sequential Memory Network (FSMN), with an appropriate number of FSMN layers and neurons per layer. The FSMN extracts features and models the sequence of the audio signals through a series of linear transformations and non-linear activation functions (such as ReLU), obtaining feature vectors. These feature vectors are then input into a fully connected layer, with the number of neurons in the fully connected layer set to the number of candidate wake words. The fully connected layer outputs the predicted probability of each candidate wake word through linear transformations and a softmax function, selecting the candidate wake word with the highest predicted probability as the identified target wake word.

[0173] A multi-channel wake word recognition model is a model used to identify the target wake word from multi-channel audio signals. For example, a feedforward sequential memory network (FSMN) can be used. By performing operations such as feature extraction, sequence modeling and linear transformation on the audio signal, the predicted probability of each candidate wake word is output, and the candidate wake word with the highest predicted probability is selected as the target wake word.

[0174] Another feasible implementation involves using Automatic Speech Recognition (ASR) technology to convert audio signals into text content. ASR typically includes steps such as feature extraction, acoustic modeling, language modeling, and decoding. First, feature extraction is performed on the audio signal, extracting MFCC features. An appropriate acoustic model type (such as a Hidden Markov Model (HMM)) is used, with suitable model parameters (such as the number of states and the transition probability matrix, estimated based on training data) set. An appropriate language model type (such as an n-gram language model) is used, with a suitable n value set. The decoding process employs the Viterbi algorithm, searching for the most probable phoneme sequence from the feature sequence using the joint probability of the acoustic and language models, and then converting the phoneme sequence into text content. Finally, by matching the multi-channel text content with pre-defined candidate wake words, the target wake word that has successfully matched is determined.

[0175] After identifying the target wake-up word, the processing flow shown in Figure 4B can be adopted, using a target sound source localization model for processing. Figure 4B is a schematic diagram of the model processing process in an embodiment of this application. The processing process implemented using the target sound source localization model will be described below with reference to Figure 4B:

[0176] Step 4021: The control device extracts the unit content vector corresponding to each pronunciation unit contained in the target wake-up word, and extracts the position prediction vector fused with audio content information and the audio content vector fused with the position information of the collected object based on the multi-channel audio signal. The position prediction vector is used to describe the spatial position corresponding to at least one collected object with speech expression.

[0177] Specifically, in step 4021, the trained target sound source localization model is used to first analyze and determine each pronunciation unit contained in the target wake word, and encode the corresponding initial vector for each pronunciation unit; then, a self-attention mechanism is used to capture the correlation between the initial vectors to obtain the unit content vector determined for each pronunciation unit.

[0178] Subsequently, based on the multi-channel audio signals, the location prediction vector fused with audio content information and the audio content vector fused with the location information of the collected objects are extracted. Based on the multi-channel audio signals, the content description vector representing the speech content and the location description vector representing the spatial location of the collected objects are extracted. Then, the content description vector and the location description vector mapped to the same vector space are superimposed to obtain the comprehensive information vector. Based on the comprehensive information vector, the location prediction vector fused with audio content information and the audio content vector fused with the location information of each collected object are extracted.

[0179] The related processing and training processes use the same descriptive logic, and will not be elaborated further in this application.

[0180] In this way, by analyzing the target wake word, text content-level processing can be performed to obtain content vectors for each unit that incorporate the influence of pronunciation units from other locations. In addition, by extracting vectors from multi-channel audio signals, audio signal-level processing can be performed, enabling the extraction of corresponding content description vectors and location description vectors from the perspectives of audio content and the location of each acquired object, respectively, based on the multi-channel audio signals. By leveraging vector superposition and vector extraction during the feature extraction process, the influence of speaker location information can be incorporated into the extracted audio content vectors, and the influence of audio content information can be incorporated into the extracted location prediction vectors. This enhances the connection between feature vectors from different perspectives and helps improve the construction effect when building location description vectors for speech units.

[0181] Step 4022: The control device adjusts the position prediction vector according to the correlation between each unit content vector and the audio content vector, obtains the sound source position indication vector of the corresponding pronunciation unit of each unit content vector, and determines the location of the sound source object of the target wake word based on the obtained sound source position indication vector.

[0182] When performing step 4022, within the target sound source localization model, the content vector of each unit can be used as the Q vector in the attention network, the determined audio content vector can be used as the K vector in the attention network, and the position prediction vector can be used as the V vector in the attention network. With the help of the processing within the attention network, the position prediction vector can be adjusted according to the degree of correlation between each unit content vector and the audio content vector, so as to obtain the sound source position indication vector of the corresponding pronunciation unit of each unit content vector.

[0183] Furthermore, with the help of a fully connected layer, the corresponding mapping positions are obtained based on the location indicator vectors of each sound source, and the location of the sound source object can be obtained based on each mapping position.

[0184] Specifically, in the process of determining the location of the sound source object of the target wake word based on the obtained sound source location indication vectors, vector mapping can be performed on each sound source location indication vector to obtain the mapped position and the corresponding prediction probability. The mapped position is the target position whose prediction probability satisfies the first preset condition after obtaining the prediction probability of each candidate position. Then, the mapped position whose total occurrence satisfies the second preset condition is used as the prediction position determined for the sound source object of the target wake word.

[0185] The first preset condition can be: the corresponding prediction probability is the highest; the second preset condition can be: the total number of occurrences is the highest; the total number of occurrences is determined by statistically analyzing the occurrence of each mapping position.

[0186] This is equivalent to breaking down the problem of locating the sound source object, starting from each speech unit expressed by the sound source object, and performing position prediction separately, which can improve the accuracy of sound source object location.

[0187] Furthermore, after determining the location of the sound source object, the control device can generate a corresponding response voice for the location of the sound source object and play the response voice to the sound source object; then, it can acquire the audio content collected for the sound source object at the location, and based on the audio content, analyze and determine the processing instructions for the sound source object and the controlled device to be controlled, and process the controlled device associated with the sound source object according to the processing instructions.

[0188] The response voice refers to the voice generated by the control device based on the location of the sound source object after the location of the target wake-up word has been determined. This voice is used to indicate that the target wake-up word triggered by the sound source object has been recognized. Playing the response voice enables an initial response to the sound source object.

[0189] Processing instructions refer to the operation instructions issued by the sound source object to the controlled device after the control device acquires the audio content to be collected from the sound source object at the positioning location and analyzes the audio content (such as text conversion processing and semantic analysis). These instructions are used to perform targeted processing on the controlled device associated with the sound source object.

[0190] Specifically, to indicate that the target wake-up word triggered by the sound source object has been identified, the control device can generate a response voice based on the location of the sound source object. Furthermore, in the subsequent response process, the audio acquisition device associated with the sound source object's location can be identified, and only the single-channel audio signal acquired by that device can be analyzed to determine the processing instructions contained within it. The associated controlled device is then processed according to these instructions. Specifically, when analyzing the processing instructions within the single-channel audio signal, text conversion processing can be performed to obtain the corresponding text results. Semantic analysis of these text results determines the targeted controlled device and processing method. Based on the processing instructions jointly represented by the controlled device and processing method, targeted processing is achieved.

[0191] In this way, after locating the sound source object, a specific response can be made to the sound source object, and targeted processing can be carried out according to the processing instructions of the sound source object to meet the personalized needs of the sound source object, realize voice interaction with the sound source object, and ensure the user experience of the sound source object.

[0192] Furthermore, in this embodiment, after processing the controlled device associated with the sound source object according to the processing instructions of the sound source object, the control device can generate a prompt voice indicating that the processing is complete and play the prompt voice to the sound source object; then, it acquires the multi-channel audio signals that the audio acquisition component continues to acquire, and performs matching processing with each candidate wake-up word based on the continuously acquired multi-channel audio signals. The prompt voice refers to the voice indicating that the processing is complete generated by the control device after processing the controlled device associated with the sound source object according to the processing instructions of the sound source object, and by playing the prompt voice to the sound source object, it informs the sound source object that the processing is complete.

[0193] It should be noted that when playing the prompt voice, it can be played using a pre-configured controllable speaker of a nearby sound source object, or, if the sound source object is wearing controllable headphones, the controllable headphones worn by the sound source object can be instructed to play the prompt voice.

[0194] In a feasible implementation of this application, the control device can modularize the sound source positioning function based on the target sound source positioning model, so that the sound source positioning function can be realized by means of the positioning module.

[0195] It should be understood that, in the embodiments of this application, during the processing of a sound source object's processing instruction, it is not necessary to receive response triggers from other objects based on wake words. In addition, before generating a prompt voice indicating that the processing is complete, it is necessary to determine whether the response to the sound source object is complete. The basis for determining whether the response to the sound source object is complete can be: determining that the sound source object has not made any speech expression within Z seconds; or, the audio signal collected by the audio acquisition device for the sound source object does not include speech content related to the controlled device requested by the sound source object. The value of Z is set according to the actual processing needs. Moreover, possible related words can be preset for various controllable devices. Based on this, by converting the audio signal into text content and determining whether the converted text includes the preset related words, it is possible to determine whether the audio signal includes speech content related to the controlled device.

[0196] For example, if the controlled device is a "car window", then the associated words configured for "car window" can be: car window, glass, open, closed.

[0197] In this way, after completing a targeted response to the current sound source object, the multi-channel audio signals collected can continue to be processed so as to continue to provide a targeted response to the audio object that expresses the wake word.

[0198] The following, with reference to the accompanying diagram, illustrates the process of locating the sound source object in an in-vehicle scenario, using the example of locating the sound source object based on a custom wake-up word expressed through voice:

[0199] Referring to Figure 5, which is a schematic diagram of the vector change process in the process of locating the sound source object in the embodiment of this application, it can be seen from the process shown in Figure 5 that for the processing of multi-channel audio signals, firstly, a two-dimensional multi-channel convolutional network is used to extract multi-channel information to obtain a mixed audio vector; then, the mixed audio vector is sent to a speech recognition encoder network to extract the content description vector in the multi-channel audio signal. At the same time, a speech feature extraction encoder network is used to extract the position description vector in the multi-channel audio signal. Then, the content description vector and the position description vector are superimposed to obtain a comprehensive information vector with position information; then, the comprehensive information vector is sent to two fully connected layers to map to the K vector (i.e., audio content vector) and V vector (i.e., position prediction vector) in the attention network, respectively.

[0200] For the text processing, after the multi-channel audio signal is passed through the multi-channel wake-up model, if the multi-channel audio signal contains a wake-up word, the model outputs the corresponding target wake-up word, thereby obtaining the text content on which the processing is based. Then, each speech unit is determined from the target wake-up word. Optionally, each speech unit can be processed by one-hot encoding and self-attention mechanism to obtain the unit content vector corresponding to each speech unit.

[0201] Next, the unit content vector of each speech unit is used as the Q vector in the attention mechanism, and attention operation is performed with the K vector (i.e., audio content vector) and V vector (i.e., position prediction vector) in the attention network. Specifically, the Q vector and K vector are multiplied to obtain the similarity weight. Then, the similarity weight is weighted with the V vector to obtain the sound source position indication vector corresponding to the speech part of each speech unit in the target wake word. Then, with the help of a fully connected layer, the directional confidence of each passenger position in the vehicle scene can be mapped based on each sound source position indication vector.

[0202] Furthermore, from the perspective of feature dimension changes, referring to Figure 6, which is a schematic diagram of feature dimension changes within the target sound source localization model in this application embodiment, it can be seen from the content shown in Figure 6 that the processed multi-channel audio signal is assumed to be in the form of: C1*H1*W1, where C1 is the number of channels. In the case of frequency domain processing, the value of C1 is twice that of the time domain signal, that is, twice the total number of audio acquisition devices; H1 is the height, which is the number of audio frames; W1 is the width, which is usually the number of discrete points; then, after processing by the two-dimensional multi-channel convolutional network, it is equivalent to putting multiple channels in one dimension, thus obtaining a vector dimension of H1*W2, where H1 is the height and W2 is the width; after processing by the speech recognition encoder network, a content description vector with dimension H2*D1 is obtained, and at the same time, after processing by the speech feature extraction encoding network, a position description vector with dimension H2*D2 is obtained.

[0203] Furthermore, with the help of fully connected layer 1, the content description vector is mapped to the H2*D3 vector space, and with the help of fully connected layer 2, the location description vector is mapped to the H2*D3 vector space. Then, the content description vector and the location description vector in the same vector space are superimposed to obtain a comprehensive information vector with dimension H2*D3. Then, fully connected layer 3 is used to map the comprehensive information vector to the H2*D4 vector space to obtain the audio content vector (i.e., the K vector in the attention network) that incorporates the location information of the collected object. At the same time, fully connected layer 4 is used to map the comprehensive information vector to the H2*D4 vector space to obtain the location prediction vector (i.e., the V vector in the attention network) that incorporates the audio content information.

[0204] In processing the target wake word, assuming that the target wake word includes 4 speech units, then with the help of the encoding network, an encoding vector of dimension 1*dk can be obtained for each speech unit. Then, with the help of the self-attention network, the encoding vector of each speech unit is processed to obtain the unit content vector of dimension 1*dv for each speech unit. Then, with the help of the fully connected layer, the text embedding composed of the unit content vectors is mapped to the Q vector in the attention network of dimension 4*D4.

[0205] Furthermore, with the help of attention network processing, based on Q vector, K vector and V vector, a sound source location indication vector with dimension 1*D4 is obtained for each speech unit.

[0206] In summary, this application proposes a sound source object localization method applicable to custom wake words. On one hand, it combines the content vectors of each unit obtained from analyzing the target wake word in text form with the feature vectors obtained from audio signal analysis, enabling accurate localization of sound source objects and applicable to sound source localization based on any wake word. On the other hand, by introducing an attention mechanism, it can automatically learn the speech range matching each pronunciation unit in the text. Moreover, after training the model with noisy training samples, it can achieve high localization accuracy in strong interference and high noise scenarios. Furthermore, during model training, simulating the impulse response of the hybrid microphone spacing can generate multi-channel data for training in real time, thus enabling the construction of suitable training samples for any microphone spacing without the need to collect a large amount of actual data. Additionally, the sound source localization model uses a front-end and back-end joint modeling approach, eliminating the need for an acoustic front-end; the optimization objective is only localization accuracy, simplifying the processing flow and achieving end-to-end modeling of the localization problem. This significantly improves the system's localization accuracy and exhibits high robustness in difficult scenarios, preventing a decrease in localization accuracy due to performance differences between separate front-ends.

[0207] In addition, through testing, the technical solution proposed in this application can achieve a positioning accuracy of over 98% in real vehicle test sets under quiet, high-noise, and high-interference scenarios.

[0208] Based on the same inventive concept, referring to Figure 7, which is a schematic diagram of the logical structure of the training device for the sound source localization model in an embodiment of this application, the training device 700 for the sound source localization model includes a training unit 701, wherein,

[0209] Training unit 701 is used to perform multiple rounds of iterative training on the constructed initial sound source localization model using various training samples. One round of iteration includes:

[0210] Read training samples; the training samples include: multi-channel sample audio signals labeled with sample wake words, and position labels corresponding to the multi-channel sample audio signals;

[0211] For each articulation unit contained in the sample wake word, the unit content vector corresponding to each articulation unit is extracted. Based on the multi-channel sample audio signal, a position prediction vector fused with audio content information and an audio content vector fused with the position information of each collected object are extracted. The position prediction vector is used to describe the spatial position of at least one collected object with speech expression.

[0212] Based on the correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of the corresponding pronunciation unit of each unit content vector. Based on the obtained sound source position indication vectors, the predicted position of the sound source object of the sample wake word is determined.

[0213] The model parameters of the sound source localization model are adjusted based on the difference between the predicted location and the corresponding location label.

[0214] Optionally, the multi-channel sample audio signal in a training sample is obtained by the analog unit 702 in the device in the following manner:

[0215] Among the multi-channel impulse responses associated with the tag locations, a target impulse response is selected; wherein, a multi-channel impulse response includes: the response result of an impulse signal sent from a tag location, which is determined by simulation for multiple preset audio acquisition locations in the sound source localization scenario; a tag location is selected from the candidate locations of the sound source object in the sound source localization scenario;

[0216] A single-channel audio signal containing a sample wake word is acquired, and the corresponding multi-channel sample audio signal is simulated based on the single-channel audio signal and the target impulse response.

[0217] Optionally, based on the single-channel audio signal and the target impulse response, the corresponding multi-channel sample audio signal is simulated, and the simulation unit 702 is used for:

[0218] The target impulse response is convolved and fused with a single-channel audio signal to obtain a multi-channel analog audio signal.

[0219] The multi-channel analog audio signal is identified as a multi-channel sample audio signal.

[0220] Optionally, after simulating the corresponding multi-channel sample audio signal, the simulation unit 702 uses any one or a combination of the following noise addition methods to process the multi-channel sample audio signal:

[0221] Multi-channel ambient noise is superimposed onto the multi-channel sample audio signal;

[0222] Construct a multi-channel interference signal with the same total number of channels as the multi-channel sample audio signal, and superimpose the multi-channel interference signal into the multi-channel sample audio signal;

[0223] Phase delay processing is performed on the multi-channel sample audio signals according to the preset phase delay levels for each of the corresponding channels.

[0224] Optionally, when configuring corresponding interference signals for each single-channel sample audio signal included in the multi-channel audio signal, the analog unit 702 performs any of the following operations:

[0225] Interference locations are selected from each label location, and the multi-channel impulse response at the interference location is obtained. Based on the multi-channel impulse response at the interference location and the single-channel audio signal that does not express the sample wake word, the corresponding multi-channel interference signal is simulated.

[0226] Based on the distribution of multiple audio acquisition locations, the corresponding signal-to-noise ratios are configured for each channel in the multi-channel sample audio signal. Based on the signal-to-noise ratios of each channel and the preset base signal, the corresponding multi-channel interference signals are obtained.

[0227] Optionally, when extracting the unit content vector corresponding to each phonetic unit contained in the sample wake word, the training unit 701 is used for:

[0228] The analysis identifies each articulation unit contained in the sample wake word, and for each articulation unit, a corresponding initial vector is encoded.

[0229] A self-attention mechanism is used to capture the correlation between the initial vectors, thereby obtaining the unit content vector determined for each phonological unit.

[0230] Optionally, when extracting a location prediction vector fused with audio content information and an audio content vector fused with the location information of each collected object based on multi-channel sample audio signals, the training unit 701 is used for:

[0231] Based on multi-channel sample audio signals, content description vectors representing speech content and location description vectors representing the spatial location of each collected object are extracted.

[0232] By superimposing the content description vector and the location description vector, which are mapped to the same vector space, a comprehensive information vector is obtained.

[0233] Based on the comprehensive information vector, a location prediction vector incorporating audio content information is extracted, as well as an audio content vector incorporating the location information of the collected object is extracted.

[0234] Optionally, when determining the predicted location of the sound source object of the sample wake word based on the obtained sound source location indication vectors, the training unit 701 is used for:

[0235] Based on the obtained sound source location indication vectors, vector mapping processing is performed to obtain the mapped position and the corresponding prediction probability. The mapped position is the target position whose prediction probability satisfies the first preset condition after obtaining the prediction probability of each candidate position.

[0236] The mapping position where the total number of occurrences satisfies the second preset condition is used as the prediction position determined for the sound source object of the sample wake word.

[0237] Based on the same inventive concept, referring to Figure 8, which is a schematic diagram of the logical structure of the sound source object positioning device in an embodiment of this application, the sound source object positioning device 800 includes an acquisition unit 801 and a positioning unit 802, wherein,

[0238] Acquisition unit 801 is used to acquire multi-channel audio signals acquired by the audio acquisition component;

[0239] The localization unit 802 is used to identify a target wake-up word that successfully matches each of the preset candidate wake-up words based on the multi-channel audio signal. It uses a trained target sound source localization model to extract the unit content vector corresponding to each pronunciation unit contained in the target wake-up word. Based on the multi-channel audio signal, it extracts the position prediction vector that integrates audio content information and the audio content vector that integrates the position information of the collected object. The position prediction vector is used to describe the spatial position of at least one collected object with speech expression. Based on the correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of each pronunciation unit corresponding to each unit content vector. Based on the obtained sound source position indication vectors, the localization position of the sound source object of the target wake-up word is determined.

[0240] Optionally, after determining the location of the sound source object of the target wake word, the device further includes a processing unit 803, which is used for:

[0241] Based on the location of the sound source object, generate the corresponding response voice and play the response voice to the sound source object;

[0242] The system acquires audio content collected from the audio source at the location, analyzes and determines the processing instructions for the audio source and the controlled device accordingly, and processes the controlled device associated with the audio source based on the processing instructions.

[0243] Optionally, after processing the controlled device associated with the sound source object, the processing unit 803 is further used to:

[0244] Generate a prompt message indicating that the representation processing is complete, and play the prompt message to the sound source object;

[0245] The system acquires multi-channel audio signals that are continuously acquired by the audio acquisition component, and performs matching processing with each candidate wake-up word based on the continuously acquired multi-channel audio signals.

[0246] Optionally, when extracting the unit content vector corresponding to each articulation unit contained in the target wake word, the localization unit 802 is used for:

[0247] The analysis identifies each articulation unit contained in the target wake word, and for each articulation unit, a corresponding initial vector is encoded.

[0248] A self-attention mechanism is used to capture the correlation between the initial vectors, thereby obtaining the unit content vector determined for each phonological unit.

[0249] Optionally, when extracting a location prediction vector fused with audio content information and an audio content vector fused with the location information of each collected object based on multi-channel audio signals, the positioning unit 802 is used for:

[0250] Based on multi-channel sample audio signals, content description vectors representing speech content and location description vectors representing the spatial location of the collected object are extracted.

[0251] The content description vector and the location description vector, which are mapped to the same vector space, are superimposed to obtain the comprehensive information vector. Based on the comprehensive information vector, the location prediction vector that incorporates audio content information is extracted, as well as the audio content vector that incorporates the location information of each collected object is extracted.

[0252] Optionally, when determining the location of the sound source object of the target wake-up word based on the obtained sound source location indication vectors, the positioning unit 802 is used to:

[0253] Based on the obtained sound source location indication vectors, vector mapping processing is performed to obtain the mapped position and the corresponding prediction probability. The mapped position is the target position whose prediction probability satisfies the first preset condition after obtaining the prediction probability of each candidate position.

[0254] The mapping position where the total number of occurrences satisfies the second preset condition is used as the prediction position determined for the sound source object of the target wake word.

[0255] For ease of description, the above sections are divided into modules (or units) according to their functions and described separately. Of course, in implementing this application, the functions of each module (or unit) can be implemented in one or more software or hardware components.

[0256] Having introduced the training method and apparatus for the sound source localization model according to exemplary embodiments of this application, as well as the sound source object localization method and apparatus, we will now introduce an electronic device according to another exemplary embodiment of this application.

[0257] Those skilled in the art will understand that the various aspects of training the sound source localization model in this application can be implemented as a system, method, or program product; and the various aspects of sound source object localization can be implemented as a system, method, or program product. Therefore, the various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which can be collectively referred to herein as a "circuit," "module," or "system."

[0258] Based on the same inventive concept as the above-described method embodiments, this application also provides an electronic device. Referring to Figure 9, which is a schematic diagram of the hardware structure of an electronic device applying an embodiment of this application, in one embodiment, the electronic device may be the control device 220 shown in Figure 2. In this embodiment, the structure of the electronic device may be as shown in Figure 9, including a memory 901, a communication module 903, and one or more processors 902.

[0259] The memory 901 is used to store computer programs executed by the processor 902. The memory 901 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0260] Memory 901 may be volatile memory, such as random-access memory (RAM); memory 901 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 901 may be any other medium capable of carrying or storing a desired computer program having the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 901 may be a combination of the above-described memories.

[0261] The processor 902 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 902 is used to implement the training method for the sound source localization model and the sound source object localization method described above when calling the computer program stored in the memory 901.

[0262] The communication module 903 is used to communicate with controlled devices and servers.

[0263] This application embodiment does not limit the specific connection medium between the memory 901, communication module 903, and processor 902. In this application embodiment, the memory 901 and processor 902 are connected via a bus 904 in Figure 9. The bus 904 is depicted as a thick line in Figure 9. The connection methods between other components are only illustrative and are not intended to be limiting. The bus 904 can be divided into address bus, data bus, control bus, etc. For ease of description, only one thick line is used to describe it in Figure 9, but it does not indicate that there is only one bus or one type of bus.

[0264] The memory 901 stores a computer storage medium containing computer-executable instructions. These instructions are used to implement the sound source localization model training method and the sound source object localization method of the embodiments of this application. The processor 902 is used to execute the aforementioned sound source localization model training method and sound source object localization method, as shown in Figures 3C and 4A-4B.

[0265] In another embodiment, the electronic device can also be other electronic devices. Referring to FIG10, which is a schematic diagram of the hardware composition structure of another electronic device applying the embodiments of this application, the electronic device can specifically be the controlled device 210 shown in FIG2. In this embodiment, the structure of the electronic device can be as shown in FIG10, including: a communication component 1010, a memory 1020, a display unit 1030, a camera 1040, a sensor 1050, an audio circuit 1060, a Bluetooth module 1070, a processor 1080, and other components.

[0266] The communication component 1010 is used to communicate with the server. In some embodiments, it may include a Circuit-Based Wireless Fidelity (WiFi) module. WiFi is a short-range wireless transmission technology, and electronic devices can use WiFi modules to help users send and receive information.

[0267] The memory 1020 can be used to store software programs and data. The processor 1080 executes various functions of the controlled device 210 and performs data processing by running the software programs or data stored in the memory 1020. In this application, the memory 1020 can store the operating system and various application programs, and, depending on actual processing needs, can also store computer programs related to the training method of the sound source localization model and the sound source object localization method of the embodiments of this application.

[0268] The display unit 1030 can also be used to display information input by the user or information provided to the user, as well as various menus of the controlled device 210, in a graphical user interface (GUI). Specifically, the display unit 1030 may include a display screen 1032 disposed on the front of the controlled device 210. The display unit 1030 can be used to display operable pages, etc.

[0269] The display unit 1030 can also be used to receive input digital or speech unit information and generate signal inputs related to user settings and function control of the controlled device 210. Specifically, the display unit 1030 may include a touch screen 1031 disposed on the front of the controlled device 210, which can collect touch operations of the user on or near it.

[0270] The touchscreen 1031 can be placed over the display screen 1032, or the touchscreen 1031 and the display screen 1032 can be integrated to realize the input and output functions of the controlled device 210. After integration, it can be referred to as a touch display screen. In this application, the display unit 1030 can display the application program and the corresponding operation steps.

[0271] Camera 1040 can be used to capture still images, which users can then post comments on via an application. An object is projected onto a photosensitive element through a lens, generating an optical image. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the processor 1080 to be converted into a digital image signal.

[0272] The controlled device may also include at least one sensor 1050, such as an accelerometer 1051, a distance sensor 1052, a fingerprint sensor 1053, and a temperature sensor 1054. The controlled device may also be equipped with other sensors such as a gyroscope, barometer, hygrometer, thermometer, infrared sensor, light sensor, and motion sensor.

[0273] Audio circuitry 1060, speaker 1061, and microphone 1062 provide an audio interface between the user and the controlled device 210. Audio circuitry 1060 converts received audio data into electrical signals and transmits them to speaker 1061, where speaker 1061 converts them into sound signals for output. Conversely, microphone 1062 converts collected sound signals into electrical signals, which are then received by audio circuitry 1060, converted back into audio data, and output to communication component 1010 for transmission to, for example, another controlled device 210, or to memory 1020 for further processing.

[0274] The Bluetooth module 1070 is used to exchange information with other Bluetooth devices that have Bluetooth modules via the Bluetooth protocol.

[0275] The processor 1080 is the control center of the controlled device, connecting various parts of the terminal via various interfaces and lines. It executes various functions of the controlled device and processes data by running or executing software programs stored in the memory 1020 and calling data stored in the memory 1020. In some embodiments, the processor 1080 may include at least one processing unit; the processor 1080 may also integrate an application processor and a baseband processor. In this application, the processor 1080 can run an operating system, applications, user interface display and touch response, as well as processing related to the training method of the sound source localization model and the sound source object localization method of the embodiments of this application. Furthermore, the processor 1080 is coupled to the display unit 1030.

[0276] In some possible implementations, the training method for the sound source localization model and the various aspects of the sound source object localization method provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product is run on an electronic device, the computer program is used to cause the electronic device to perform the steps in the training method for the sound source localization model and the sound source object localization method according to the various exemplary embodiments of this application described above. For example, the electronic device can perform the steps shown in FIG3C, 4A-4B.

[0277] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0278] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on an electronic device. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with a command execution system, apparatus, or device.

[0279] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.

[0280] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0281] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The computer program can execute entirely on the user's electronic device, partially on the user's electronic device, as a standalone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving remote electronic devices, the remote electronic device can be connected to the user's electronic device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external electronic device (e.g., via the Internet using an Internet service provider).

[0282] In summary, this application provides a training method for a sound source localization model, a sound source object localization method, a training device for a sound source localization model, a sound source object localization device, an electronic device, a computer-readable storage medium, and a computer program product. The electronic device uses various training samples to perform multiple rounds of iterative training on the constructed initial sound source localization model. In one round of iteration, the electronic device first reads training samples containing multi-channel sample audio signals labeled with sample wake words and corresponding position labels. This step provides basic data for subsequent model training, enabling the model to learn the relationship between sample wake words, multi-channel sample audio signals, and positions. Next, for each articulation unit contained in the sample wake word, unit content vectors are extracted respectively. Simultaneously, based on the multi-channel sample audio signals, position prediction vectors fused with audio content information and audio content vectors fused with the position information of each collected object are extracted. From a technical perspective, extracting unit content vectors from articulation units can refine speech information and capture the features of each articulation unit, while extracting position prediction vectors and audio content vectors comprehensively considers the audio content and the position information of the collected object, providing a more comprehensive basis for subsequent localization. Then, based on the correlation between the content vector of each unit and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of each articulatory unit, thereby determining the predicted position of the sound source object of the sample wake word. This correlation-based adjustment method can more accurately determine the sound source position according to the correlation between the articulatory unit and the audio content, avoiding the inaccurate positioning problem caused by incomplete speech separation or information loss in traditional methods. Finally, the model parameters of the sound source localization model are adjusted according to the difference between the predicted position and the position label. By continuously adjusting the parameters, the model can be gradually optimized, improving the accuracy of sound source object localization. This approach transforms the sound source object localization problem into position mapping of each speech unit, realizing end-to-end modeling of the localization process, and eliminating the need for position separation and noise reduction of multi-channel sample audio data. This simplifies the processing logic, reduces information loss and interference that may be caused by speech separation and noise reduction, improves model training effect and localization accuracy, and reduces processing resource consumption, thereby improving resource utilization.

[0283] Furthermore, the multi-channel audio signal in a training sample is obtained by selecting a target impulse response from the multi-channel impulse responses associated with each label position, acquiring a single-channel audio signal expressing the sample wake word, and simulating based on both. From a technical perspective, the multi-channel impulse response reflects the response of each audio acquisition location to signals emitted from different candidate locations. By selecting a target impulse response and combining it with the single-channel audio signal, the reception of a single-channel audio signal emitted from a candidate location by each audio acquisition location can be simulated. This simulation method makes the construction of training samples more flexible, not limited by the number and location of actual audio acquisition devices. It allows for the construction of suitable training samples according to different application scenarios and needs, improving the efficiency of training sample construction and enhancing the model's adaptability to different scenarios.

[0284] Furthermore, when simulating multi-channel sample audio signals, the target impulse response is convolved and fused with the single-channel audio signal to obtain a multi-channel simulated audio signal, which is then identified as the multi-channel sample audio signal. During the convolution fusion process, the characteristics of the target impulse response interact with the single-channel audio signal, accurately simulating physical phenomena such as attenuation and reflection during audio propagation, thus obtaining a more realistic multi-channel simulated audio signal. This directly yields a multi-channel time-domain audio signal, which can be processed into a suitable multi-channel sample audio signal according to actual needs. Both time-domain and frequency-domain signals can be easily acquired, facilitating flexible construction of training samples and providing diverse data formats for model training, helping the model learn more comprehensive audio features.

[0285] Furthermore, after simulating the multi-channel sample audio signals, noise is added using any one or a combination of methods, such as superimposing multi-channel ambient noise, superimposing multi-channel interference signals, or performing phase delay processing. In real-world audio acquisition environments, various noises and interferences, as well as phase differences between different audio acquisition devices, are unavoidable. By superimposing multi-channel ambient noise, the model can learn audio features in real-world noise environments, improving its noise resistance. Superimposing multi-channel interference signals can simulate speech interference from other locations, enabling the model to better identify the target sound source signal. Phase delay processing takes into account the quality differences between different audio acquisition devices, enhancing the model's adaptability to phase changes. Adding perturbations based on potential interference in real-world scenarios allows the training samples to simulate interference in real-world environments, enhancing the model's robustness and helping it learn to locate sound sources in high-noise and strong-interference environments, thus improving the model's localization accuracy in complex environments.

[0286] Furthermore, when constructing multi-channel interference signals, interference locations can be selected from each label position and simulated based on their multi-channel impulse response and single-channel audio signals without expressed wake words. Alternatively, the signal-to-noise ratio (SNR) can be configured for each channel of the multi-channel sample audio signal according to the distribution of multiple audio acquisition locations, combined with a preset base signal. The first method simulates speech interference from other locations, allowing the model to learn to distinguish between the target sound source and the interference source during training, thus improving the model's anti-interference ability. The second method configures the SNR based on the distribution of audio acquisition locations, considering the signal strength differences at different locations, enabling the model to better adapt to the actual audio acquisition environment. Both methods can better simulate interference in the actual audio acquisition process, making the training samples more closely match the signal state in the actual scenario, improving the construction effect of training samples, and allowing the model to learn more realistic audio features during training, further improving the model's localization accuracy and robustness.

[0287] Furthermore, when extracting the unit content vectors of each articulatory unit, the articulatory units contained in the sample wake word are first analyzed and identified, and initial vectors are encoded separately. Then, a self-attention mechanism is used to capture the correlation between the initial vectors. The self-attention mechanism can automatically learn the correlation between articulatory units and capture the contextual information in the speech. In speech signals, there may be semantic and phonological associations between different articulatory units. Through the self-attention mechanism, this association information can be integrated into the unit content vector. The resulting unit content vector better reflects the importance and correlation of the articulatory unit in the entire speech, which helps to more accurately determine the sound source location indicator vector of the articulatory unit in the subsequent process. With the help of the self-attention mechanism, the influence of articulatory units in other positions can be integrated, making the obtained unit content vector more in line with the actual use needs, improving the extraction effect of the unit content vector, and thus providing more accurate feature information for the localization of the sound source object.

[0288] Furthermore, when extracting the location prediction vector fused with audio content information and the audio content vector fused with the location information of each collected object, the process first involves extracting a content description vector representing the speech content and a location description vector representing the spatial location of each collected object based on the multi-channel sample audio signal. These two vectors are then superimposed and mapped to the same vector space to obtain a comprehensive information vector, which is then used for extraction. From a technical perspective, extracting the content description vector and location description vector separately allows for the separation and extraction of audio content and location information, facilitating subsequent processing. Superimposing and mapping them to the same vector space to obtain a comprehensive information vector allows for the fusion of speech content and the location information of the collected objects, enabling the model to consider both factors simultaneously. In the subsequent localization process, the vector fused with audio content and location information provides more comprehensive information, helping to more accurately determine the location of the sound source object. This approach integrates speaker location information into the audio content vector and audio content information into the location prediction vector, strengthening the connection between feature vectors from different angles and improving the effectiveness of subsequent location description vector construction, thereby enhancing the accuracy of sound source object localization.

[0289] Furthermore, when determining the predicted location of the sound source object of the sample wake word, vector mapping is performed based on each sound source location indicator vector to obtain the mapped location and its corresponding predicted probability. Mapped locations whose predicted probabilities satisfy a first preset condition are selected, and mapped locations whose total occurrences satisfy a second preset condition are used as the predicted locations. Vector mapping converts the sound source location indicator vector into predicted probabilities at each candidate location. By selecting mapped locations that meet the conditions, some less likely locations can be eliminated, improving the accuracy of localization. Using mapped locations whose total occurrences satisfy the conditions as the predicted locations allows for comprehensive consideration of the localization results of multiple articulatory units, avoiding the influence of localization errors from a single articulatory unit. Decomposing the sound source object localization problem and performing location prediction separately for each articulatory unit improves the accuracy of sound source object localization and also makes the localization process more flexible and reliable.

[0290] Furthermore, when the control device executes the sound source object localization method, it first acquires multi-channel audio signals collected by the audio acquisition component. Based on these signals, when a target wake-up word successfully matches one of the preset candidate wake-up words is identified, a trained target sound source localization model is used to extract unit content vectors for each articulator contained in the target wake-up word. Based on the multi-channel audio signals, a position prediction vector incorporating audio content information and an audio content vector incorporating the location information of the acquired object are extracted. The position prediction vector is adjusted based on the correlation between each unit content vector and the audio content vector to obtain the sound source position indication vector for each articulator, thereby determining the location of the sound source object of the target wake-up word. This method triggers localization processing at appropriate times, initiating the localization process only when the target wake-up word is identified, avoiding unnecessary computation and resource consumption. Predicting the location of the sound source object from the perspective of each articulator can fully utilize the detailed information in the speech, improving the accuracy of localization. Simultaneously, with the help of the trained target sound source localization model, which learns a large amount of audio features and localization information during training, it has excellent processing performance, reducing the amount of resources consumed during sound source object localization and improving resource utilization efficiency.

[0291] Furthermore, after determining the location of the sound source object corresponding to the target wake-up word, a corresponding response voice is generated and played at that location. The system then acquires further audio content, analyzes it to determine the processing instructions for the sound source object and the corresponding controlled device, and performs the processing accordingly. Generating a response voice promptly informs the sound source object that its wake-up word has been recognized, improving the user's interactive experience. After acquiring further audio content, analyzing and determining the processing instructions and controlled device allows for targeted processing based on the sound source object's needs. This approach enables specific responses to the sound source object after location, providing targeted processing based on its instructions to meet its personalized needs, achieving voice interaction, ensuring a good user experience for the sound source object, and improving the system's intelligence and practicality.

[0292] Furthermore, after processing the controlled device associated with the sound source object, a prompt voice indicating the completion of processing is generated and played. The audio acquisition component continues to acquire multi-channel audio signals and performs matching processing with each candidate wake-up word. Generating the prompt voice promptly informs the sound source object that processing is complete, allowing the user to understand the system's working status. Continuing the matching processing with each candidate wake-up word allows for continuous monitoring of new wake-up words, enabling a response to audio objects expressing new wake-up words. This allows for continued processing of the acquired multi-channel audio signals after completing a targeted response to the current sound source object, maintaining continuous system operation and responsiveness, and improving system efficiency and usability.

[0293] Furthermore, when extracting the unit content vectors of each articulation unit of the target wake word, each articulation unit is first analyzed and identified, and initial vectors are encoded separately. Then, a self-attention mechanism is used to capture relevance. Similar to the processing method in the training process, the self-attention mechanism can capture the relevance between articulation units, integrating contextual information into the unit content vectors. During localization, accurate unit content vectors provide a more reliable basis for determining the sound source location indicator vector of the articulation unit, thereby improving localization accuracy. Integrating the influence of articulation units in other positions makes the unit content vectors more consistent with actual needs, improving the accuracy of articulation unit feature extraction during localization and further ensuring the accuracy of sound source object localization.

[0294] Furthermore, when extracting the location prediction vector fused with audio content information and the audio content vector fused with the location information of each collected object based on the multi-channel audio signal, the content description vector representing the speech content and the location description vector representing the spatial location of the collected object are first extracted. These two vectors, mapped to the same vector space, are then superimposed to obtain a comprehensive information vector, which is then extracted. This processing method can also fuse audio content information and location information during the localization process, providing more comprehensive information for localization. It enhances the connection between feature vectors from different angles, helping to improve the effectiveness of constructing location description vectors for speech units during localization, thereby improving the accuracy of sound source object localization and enabling the system to more accurately determine the location of the sound source object.

[0295] Furthermore, when determining the location of the sound source object of the target wake word, vector mapping is performed based on the location indicator vectors of each sound source to obtain the mapped position and its corresponding predicted probability. Mapped positions that meet the criteria are then selected, and the total number of mapped positions meeting the criteria is taken as the location. Consistent with the processing method in the training process, this approach breaks down the localization problem, performing position prediction separately for each speech unit. This fully utilizes the information from each speech unit, avoids the influence of localization errors in individual speech units, and improves the accuracy of sound source object localization. Through vector mapping and selection, the location of the sound source object can be determined more accurately, providing a reliable foundation for subsequent targeted responses and processing.

[0296] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0297] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0298] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing a computer-usable computer program.

[0299] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0300] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0301] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for training a sound source localization model, executed by an electronic device, using various training samples to perform multiple rounds of iterative training on the constructed initial sound source localization model, wherein, One round of iteration includes: Read training samples; the training samples include: multi-channel sample audio signals labeled with sample wake words, and position labels corresponding to the multi-channel sample audio signals; For each articulation unit contained in the sample wake word, the unit content vector corresponding to each articulation unit is extracted. Based on the multi-channel sample audio signal, a position prediction vector fused with audio content information and an audio content vector fused with the position information of each collected object are extracted. The position prediction vector is used to describe the spatial position of at least one collected object with speech expression. Based on the correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of the corresponding pronunciation unit for each unit content vector. Based on the obtained sound source position indication vectors, the predicted position of the sound source object of the sample wake word is determined. Based on the difference between the predicted location and the corresponding location label, the model parameters of the sound source localization model are adjusted.

2. The method as described in claim 1, wherein the multi-channel sample audio signal in a training sample is obtained in the following manner: Among the multi-channel impulse responses associated with the respective label locations, a target impulse response is selected; where, A multi-channel impulse response includes: a response result for an impulse signal sent from a tag location, which is determined by simulation for multiple preset audio acquisition locations in the sound source localization scenario; a tag location is selected from candidate locations of the sound source object in the sound source localization scenario. A single-channel audio signal containing a sample wake word is acquired, and a corresponding multi-channel sample audio signal is simulated based on the single-channel audio signal and the target impulse response.

3. The method as described in claim 2, wherein simulating the corresponding multi-channel sample audio signal based on the single-channel audio signal and the target impulse response comprises: The target impulse response is convolved and fused with the single-channel audio signal to obtain a multi-channel analog audio signal; The multi-channel analog audio signal is determined to be a multi-channel sample audio signal.

4. The method as described in claim 2, wherein after simulating the corresponding multi-channel sample audio signal, the multi-channel sample audio signal is processed using any one or a combination of the following noise addition methods: Multi-channel environmental noise is superimposed on the multi-channel sample audio signal; Construct a multi-channel interference signal with the same total number of channels as the multi-channel sample audio signal, and superimpose the multi-channel interference signal onto the multi-channel sample audio signal; The multi-channel sample audio signals are subjected to phase delay processing according to the preset phase delay levels for each of the corresponding channels.

5. The method of claim 4, wherein constructing a multi-channel interference signal with the same total number of channels as the multi-channel sample audio signal comprises any one of the following: An interference location is selected from each label location, and the multi-channel impulse response at the interference location is obtained. Based on the multi-channel impulse response at the interference location and the single-channel audio signal that does not express the sample wake word, the corresponding multi-channel interference signal is simulated. Based on the distribution of the multiple audio acquisition locations, a corresponding signal-to-noise ratio is configured for each channel in the multi-channel sample audio signal. Then, based on the signal-to-noise ratio of each channel and a preset base signal, a corresponding multi-channel interference signal is obtained.

6. The method according to any one of claims 1-5, wherein extracting the unit content vector corresponding to each pronunciation unit contained in the sample wake word includes: The analysis identifies each articulation unit contained in the sample wake word, and for each articulation unit, a corresponding initial vector is encoded. A self-attention mechanism is used to capture the correlation between the initial vectors to obtain the unit content vectors determined for each of the articulation units.

7. The method according to any one of claims 1-5, wherein extracting a location prediction vector fused with audio content information and an audio content vector fused with the location information of each collected object based on the multi-channel sample audio signal comprises: Based on the multi-channel sample audio signals, content description vectors representing speech content and location description vectors representing the spatial location of each collected object are extracted. The content description vector and the position description vector, which are mapped to the same vector space, are superimposed to obtain a comprehensive information vector; Based on the comprehensive information vector, a location prediction vector incorporating audio content information is extracted, and an audio content vector incorporating the location information of the collected object is extracted.

8. The method according to any one of claims 1-5, wherein determining the predicted location of the sound source object of the sample wake word based on the obtained sound source location indication vectors comprises: Based on the obtained sound source position indication vectors, vector mapping processing is performed to obtain the mapped position and the corresponding prediction probability. The mapped position is the target position whose prediction probability satisfies the first preset condition after obtaining the prediction probability of each candidate position. The mapping position where the total number of occurrences satisfies the second preset condition is used as the prediction position determined for the sound source object of the sample wake word.

9. A method for locating a sound source object, executed by a control device, comprising: Acquire multi-channel audio signals from the audio acquisition component; Based on the multi-channel audio signal, when identifying a target wake-up word that successfully matches each of the preset candidate wake-up words, a trained target sound source localization model is used to extract the unit content vector corresponding to each articulation unit contained in the target wake-up word. Furthermore, based on the multi-channel audio signal, a position prediction vector incorporating audio content information and an audio content vector incorporating the position information of the captured object are extracted. The position prediction vector describes the spatial position of at least one captured object with a spoken expression. Based on the correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of the pronunciation unit corresponding to each unit content vector, and the location of the sound source object of the target wake word is determined based on the obtained sound source position indication vector.

10. The method of claim 9, further comprising, after determining the location of the sound source object of the target wake-up word: For the location of the sound source object, generate a corresponding response voice and play the response voice to the sound source object; The system acquires audio content collected from the audio source object at the specified location, analyzes and determines the processing instructions for the audio source object and the corresponding controlled device based on the audio content, and processes the controlled device associated with the audio source object according to the processing instructions.

11. The method of claim 10, wherein after processing the controlled device associated with the sound source object, it further includes: Generate a prompt voice indicating completion of representation processing, and play the prompt voice to the sound source object; The multi-channel audio signals continuously acquired by the audio acquisition component are obtained, and matching processing with each candidate wake-up word is performed based on the continuously acquired multi-channel audio signals.

12. The method according to any one of claims 9-11, wherein extracting the unit content vector corresponding to each pronunciation unit contained in the target wake word includes: The analysis identifies each articulation unit contained in the target wake word, and for each articulation unit, a corresponding initial vector is encoded. A self-attention mechanism is used to capture the correlation between the initial vectors to obtain the unit content vectors determined for each of the articulation units.

13. The method according to any one of claims 9-11, wherein extracting a location prediction vector fused with audio content information and an audio content vector fused with the location information of each collected object based on the multi-channel audio signal comprises: Based on the multi-channel audio signal, a content description vector representing the speech content and a location description vector representing the spatial location of the collected object are extracted. The content description vector and the location description vector, which are mapped to the same vector space, are superimposed to obtain a comprehensive information vector. Based on the comprehensive information vector, a location prediction vector that incorporates audio content information is extracted, and an audio content vector that incorporates the location information of each collected object is extracted.

14. The method according to any one of claims 9-11, wherein determining the location of the sound source object of the target wake-up word based on the obtained sound source location indication vectors comprises: Based on the obtained sound source position indication vectors, vector mapping processing is performed to obtain the mapped position and the corresponding prediction probability. The mapped position is the target position whose prediction probability satisfies the first preset condition after obtaining the prediction probability of each candidate position. The mapping position where the total number of occurrences satisfies the second preset condition is used as the prediction position determined for the sound source object of the target wake word.

15. A training device for a sound source localization model, the device comprising: The training unit is used to perform multiple rounds of iterative training on the constructed initial sound source localization model using various training samples. One round of iteration includes: Read training samples; the training samples include: multi-channel sample audio signals labeled with sample wake words, and position labels corresponding to the multi-channel sample audio signals; For each articulation unit contained in the sample wake word, the unit content vector corresponding to each articulation unit is extracted. Based on the multi-channel sample audio signal, a position prediction vector fused with audio content information and an audio content vector fused with the position information of each collected object are extracted. The position prediction vector is used to describe the spatial position of at least one collected object with speech expression. Based on the degree of correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted respectively to obtain the sound source position indication vector of the pronunciation unit corresponding to each unit content vector, and based on the obtained sound source position indication vector, the predicted position of the sound source object of the sample wake word is determined. Based on the difference between the predicted location and the corresponding location label, the model parameters of the sound source localization model are adjusted.

16. A sound source object positioning device, comprising: The acquisition unit is used to acquire multi-channel audio signals collected by the audio acquisition component; The localization unit, when identifying a target wake-up word that successfully matches each of the preset candidate wake-up words based on the multi-channel audio signal, uses a trained target sound source localization model to extract the unit content vector corresponding to each pronunciation unit contained in the target wake-up word, and extracts a position prediction vector fused with audio content information and an audio content vector fused with the position information of the collected object based on the multi-channel audio signal. The position prediction vector is used to describe the spatial position corresponding to at least one collected object with speech expression. Based on the correlation between each unit content vector and the audio content vector, the position prediction vector is adjusted to obtain the sound source position indication vector of the pronunciation unit corresponding to each unit content vector, and the localization position of the sound source object of the target wake-up word is determined based on the obtained sound source position indication vectors.

17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1-14.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method as described in any one of claims 1-14.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method as claimed in any one of claims 1-14.

Citation Information

Patent Citations

  • Sound source positioning method and device, equipment and computer storage medium

    CN112201259A

  • Sound source positioning method and device, computer readable storage medium and electronic equipment

    CN112799016A

  • Sound source localization model training and sound source localization method and device

    CN113903334A

  • Sound source positioning model training method, sound source object positioning method and related device

    CN118675507A

  • Microphone position determination device and microphone position determination method

    JP2019097100A