Signal processing method, model training method, device, equipment, medium and product

By processing multi-channel audio signals through an end-to-end multi-task neural network model, fusion features are obtained and the direction of the sound source is located. This solves the problem of voice interaction error accumulation in complex environments of electronic devices and improves the accuracy of voice interaction.

CN120998181APending Publication Date: 2025-11-21MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511285395.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

When electronic devices perform voice interaction in complex environments, they are affected by location and sound interference, which leads to the accumulation of errors and affects the accuracy of voice interaction response.

Method used

An end-to-end multi-task neural network model is adopted. The feature acquisition module obtains the fusion features of multi-channel audio signals, the sound source localization module obtains the sound source direction features, and the shared coding module extracts the context information and provides it to the speech task module for processing.

Benefits of technology

It improves the focusing accuracy and anti-interference ability of target speech in complex environments, overcomes the cumulative error of cascaded methods, and enhances the accuracy of voice interaction response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998181A_ABST
    Figure CN120998181A_ABST
Patent Text Reader

Abstract

The invention provides a signal processing method and device, a model training method and device, equipment, a medium and a product. The method comprises the following steps: acquiring audio signals of multiple channels; a feature acquisition module is called, a first fusion feature is acquired according to the audio signals of the multiple channels, and the first fusion feature is used for representing text information and spatial information of the audio signals of the multiple channels; calling a sound source positioning module, and obtaining a first sound source direction feature according to the first fusion feature; calling a feature fusion module, and fusing the first fusion feature and the first sound source direction feature to obtain a second fusion feature; and calling a shared coding module, performing context information extraction on the second fusion feature to obtain a first shared feature, the first shared feature being provided for each voice task module in the at least one voice task module to execute a corresponding voice task. Based on the scheme of the invention, the accuracy of voice interaction response of the electronic equipment can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio signal processing, and in particular to a signal processing method, a model training method, an apparatus, a device, a medium and a product. BACKGROUND

[0002] Currently, some electronic devices can collect audio signals containing user speech, and then use the audio signals to realize voice interaction functions. In actual application scenarios, due to the influence of the location of the user or other sound interference, errors may be generated in the process of processing the audio signals by the electronic device, and these errors may be accumulated, ultimately affecting the accuracy of the voice interaction response of the electronic device. SUMMARY

[0003] The present application provides a signal processing method, a model training method, an apparatus, a device, a medium and a product, which effectively improve the accuracy of the voice interaction response of the electronic device.

[0004] In a first aspect, a signal processing method is provided, applied to an electronic device, the electronic device being provided with a target model, the target model comprising a feature acquisition module, a sound source positioning module, a feature fusion module, a shared coding module and at least one speech task module, and the method comprising:

[0005] obtaining audio signals of multiple channels;

[0006] calling the feature acquisition module to obtain first fusion features from the audio signals of the multiple channels, the first fusion features being used to represent text information and spatial information of the audio signals of the multiple channels;

[0007] calling the sound source positioning module to obtain first sound source direction features from the first fusion features;

[0008] calling the feature fusion module to fuse the first fusion features and the first sound source direction features to obtain second fusion features;

[0009] calling the shared coding module to extract context information from the second fusion features to obtain first shared features, the first shared features being used to provide each speech task module in the at least one speech task module to execute a corresponding speech task.

[0010] In combination with the first aspect, in some possible implementation manners, calling the feature acquisition module to obtain the first fusion features from the audio signals of the multiple channels comprises: calling the feature acquisition module to extract target features corresponding to the audio signals of each channel in the audio signals of the multiple channels, the target features being used to represent text information and spatial information of the audio signals of the corresponding channel; and fusing the target features corresponding to the audio signals of each channel to obtain the first fusion features.

[0011] With reference to the first aspect and the foregoing implementation manners, in some possible implementation manners, the calling the sound source positioning module and obtaining the first sound source direction feature according to the first fused feature comprises: calling the sound source positioning module to perform direction classification processing on the first fused feature to obtain a direction classification result; and generating a corresponding direction embedding vector according to the direction classification result; and determining the direction embedding vector as the first sound source direction feature.

[0012] With reference to the first aspect and the foregoing implementation manners, in some possible implementation manners, the at least one speech task module comprises a speech wake-up module; after the calling the feature obtaining module and obtaining the first fused feature according to the audio signals of the plurality of channels, and before the calling the sound source positioning module and obtaining the first sound source direction feature according to the first fused feature, the method further comprises: calling the feature fusion module to fuse the first fused feature with a preset second sound source direction feature to obtain a third fused feature; calling the shared coding module to perform context information extraction on the third fused feature to obtain a second shared feature; calling the speech wake-up module to perform wake-up instruction analysis according to the second shared feature to obtain a wake-up instruction analysis result; and if the electronic device is controlled to enter a wake-up state according to the wake-up instruction analysis result, performing the step of calling the sound source positioning module and obtaining the first sound source direction feature according to the first fused feature.

[0013] With reference to the first aspect and the foregoing implementation manners, in some possible implementation manners, the at least one speech task module comprises a speech control module; after the calling the shared coding module and performing context information extraction on the second fused feature to obtain the first shared feature, the method further comprises: calling the speech control module to perform control instruction analysis according to the first shared feature to obtain a control instruction analysis result; and if a speech control instruction is generated according to the control instruction analysis result, controlling the electronic device to perform a corresponding speech control function according to the speech control instruction.

[0014] With reference to the first aspect and the foregoing implementation manners, in some possible implementation manners, the at least one speech task module comprises a speech enhancement module; after the calling the shared coding module and performing context information extraction on the second fused feature to obtain the first shared feature, the method further comprises: calling the speech enhancement module to obtain an enhanced speech signal according to the first shared feature; and sending the enhanced speech signal to a server in communication connection with the electronic device, so that the server interacts with the electronic device according to the enhanced speech signal.

[0015] In a second aspect, a model training method is provided, which comprises:

[0016] obtaining an initial model, the initial model comprising a feature obtaining module, a sound source positioning module, a feature fusion module, a shared coding module, and at least one speech task module;

[0017] obtain a training sample, the training sample comprising audio signals of multiple channels, labeled information of a sound source positioning module, and labeled information of each speech task module in at least one speech task module;

[0018] invoke a feature obtaining module to obtain first fusion features corresponding to the training sample according to the audio signals of the multiple channels in the training sample, the first fusion features corresponding to the training sample being used to represent text information and spatial information of the audio signals of the multiple channels;

[0019] invoke the sound source positioning module to obtain first sound source direction features corresponding to the training sample according to the first fusion features corresponding to the training sample;

[0020] invoke a feature fusion module to fuse the first fusion features corresponding to the training sample and the first sound source direction features to obtain second fusion features corresponding to the training sample;

[0021] invoke a shared encoding module to extract context information from the second fusion features corresponding to the training sample to obtain first shared features corresponding to the training sample;

[0022] update the initial model based on the first shared features corresponding to the training sample, the labeled information of the sound source positioning module, and the labeled information of each speech task module to obtain a target model.

[0023] With reference to the second aspect, in some possible implementation manners, invoking the feature obtaining module to obtain the first fusion features corresponding to the training sample according to the audio signals of the multiple channels in the training sample comprises: invoking the feature obtaining module to extract target features corresponding to the audio signals of each channel in the training sample, the target features being used to represent text information and spatial information of the audio signals of the corresponding channel; and fusing the target features corresponding to the audio signals of each channel in the training sample to obtain the first fusion features corresponding to the training sample.

[0024] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, invoking the sound source positioning module to obtain the first sound source direction features corresponding to the training sample according to the first fusion features corresponding to the training sample comprises: invoking the sound source positioning module to perform direction classification processing on the first fusion features corresponding to the training sample to obtain direction classification results corresponding to the training sample; generating a direction embedding vector corresponding to the training sample according to the direction classification results corresponding to the training sample; and determining the direction embedding vector corresponding to the training sample as the first sound source direction features corresponding to the training sample.

[0025] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, based on the first shared feature corresponding to the training sample, the labeling information of the sound source positioning module, and the labeling information of each speech task module, the initial model is updated in parameters to obtain the target model, including: determining a loss parameter of the sound source positioning module under the training sample according to the direction classification result corresponding to the training sample and the labeling information of the sound source positioning module in the training sample, the direction classification result corresponding to the training sample being obtained in the process of obtaining the first sound source direction feature corresponding to the training sample; calling each speech task module, and executing a speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each speech task module under the training sample; determining a loss parameter of each speech task module under the training sample according to the task execution result of each speech task module under the training sample and the labeling information of each speech task module in the training sample; constructing a joint loss parameter corresponding to the training sample according to the loss parameter of the sound source positioning module under the training sample and the loss parameters of each speech task module under the training sample; and updating the initial model in parameters according to the joint loss parameter corresponding to the training sample to obtain the target model.

[0026] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the at least one speech task module includes a speech wake-up module; the calling of each speech task module and the execution of a speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each speech task module under the training sample includes: calling the speech wake-up module, and performing wake-up instruction analysis according to the first shared feature corresponding to the training sample to obtain a wake-up instruction analysis result of the speech wake-up module under the training sample; and determining the wake-up instruction analysis result of the speech wake-up module under the training sample as the task execution result of the speech wake-up module under the training sample.

[0027] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the at least one speech task module includes a speech control module; the calling of each speech task module and the execution of a speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each speech task module under the training sample includes: calling the speech control module, and performing control instruction analysis according to the first shared feature corresponding to the training sample to obtain a control instruction analysis result of the speech control module under the training sample; and determining the control instruction analysis result of the speech control module under the training sample as the task execution result of the speech control module under the training sample.

[0028] With reference to the second aspect and the foregoing implementation manners, in a possible implementation manner, the at least one speech task module includes a speech enhancement module; and the calling each speech task module to perform a speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each speech task module under the training sample includes: calling the speech enhancement module to obtain an enhanced speech signal corresponding to the training sample according to the first shared feature corresponding to the training sample; and determining the enhanced speech signal corresponding to the training sample as the task execution result of the speech enhancement module under the training sample.

[0029] In a third aspect, an electronic device is provided, and the electronic device includes:

[0030] a memory configured to store executable program code;

[0031] a processor configured to call and run the executable program code from the memory, so that the electronic device performs the method of any one of the foregoing aspects.

[0032] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program. When the computer program is executed, the method of any one of the foregoing aspects is implemented.

[0033] The technical solutions provided by some embodiments of the present application have at least the following beneficial effects. First, the audio signals of multiple channels are obtained, providing a data basis for subsequent processing. Then, the feature acquisition module is called to obtain first fusion features from the audio signals of multiple channels, to implement preliminary integration of the text information and spatial information of the audio signals. Then, the sound source positioning module is called to obtain first sound source direction features from the first fusion features, and the first sound source direction features are used as spatial prior information to enhance the spatial perception capability. Then, the feature fusion module is called to fuse the first fusion features and the first sound source direction features to obtain second fusion features. Finally, the shared encoding module is called to extract context information from the second fusion features to obtain first shared features. The first shared features contain spatial prior information, as well as text information and spatial information of the audio signals, which can be provided to each speech task module to perform a corresponding speech task. In this way, on the one hand, the spatial prior information is integrated into the first shared features, improving the focusing precision and anti-interference capability of the target speech in a complex environment. On the other hand, through end-to-end multi-module collaborative processing, the cumulative error that may be caused by the cascading manner is overcome, and finally the accuracy of the speech interaction response of the electronic device can be effectively improved. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the accompanying drawings in the following description only only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0035] Figure 1 is a scene schematic diagram of a signal processing provided by an embodiment of the present application;

[0036] Figure 2 is a flow schematic diagram of a signal processing method provided by an embodiment of the present application;

[0037] Figure 3 is a flow schematic diagram of obtaining a first fusion feature provided by an embodiment of the present application;

[0038] Figure 4 is a flow schematic diagram of obtaining a first sound source direction feature provided by an embodiment of the present application;

[0039] Figure 5 is a flow schematic diagram of introducing a sound source direction after wake-up provided by an embodiment of the present application;

[0040] Figure 6 is an example schematic diagram of model inference provided by an embodiment of the present application;

[0041] Figure 7 is a flow schematic diagram of a model training method provided by an embodiment of the present application;

[0042] Figure 8 is a flow schematic diagram of obtaining a first fusion feature corresponding to a training sample provided by an embodiment of the present application;

[0043] Figure 9 is a flow schematic diagram of obtaining a first sound source direction feature corresponding to a training sample provided by an embodiment of the present application;

[0044] Figure 10 is a flow schematic diagram of training a target model provided by an embodiment of the present application;

[0045] Figure 11 is an example schematic diagram of model training provided by an embodiment of the present application;

[0046] Figure 12 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make the features and advantages of the present application more apparent, the following will describe the technical solutions in the embodiments of the present application in a clear and complete manner with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative efforts are within the scope of the present application.

[0048] The following description refers to the accompanying drawings. Unless otherwise indicated, same or similar elements in different drawings have same or similar reference numerals. The implementations described in the following exemplary embodiments are not meant to represent all implementations consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.

[0049] Hereinafter, the terms "first" and "second" are used only for the purpose of description, and should not be understood as implying or suggesting relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" can explicitly or implicitly include one or more of the features.

[0050] The following will be described in detail respectively. It should be noted that the sequence of the following embodiment descriptions does not limit the preferred sequence of the embodiments.

[0051] Please refer to Figure 1 , Figure 1 is a scene schematic diagram of signal processing provided by an embodiment of the present application. For some electronic devices supporting voice interaction function, a user can issue a voice near the electronic devices, and the sound wave is propagated through air medium. The electronic devices collect multi-channel audio signals through built-in or external microphone arrays. The audio signals contain text information of the user voice, and then the voice interaction function can be realized by using the audio signals. Exemplarily, the electronic devices can be air conditioning devices, air supply devices, display devices, etc. Taking the air conditioning device as an example, the user can control temperature adjustment, mode switching or air speed adjustment through voice; taking the air supply device as an example, the user can control the device to start or stop or swing the angle through voice; taking the display device as an example, the user can switch channels, adjust the volume or start the application through voice. In addition to the above-mentioned air conditioning device, air supply device and display device, the electronic device can also be other devices supporting voice interaction function, which will not be listed one by one here.

[0052] To process the aforementioned audio signals, the relevant technology employs multiple independently trained models, cascading them to sequentially perform feature extraction, voice wake-up, and voice control. The input of the next model depends on the output of the previous model. In real-world applications, users may speak from anywhere in the space where the electronic device is located, or the space may be filled with various types of sound interference, potentially leading to errors in the output results of multiple models. Furthermore, since the multiple models are independently trained, the output accuracy of some models is difficult to guarantee due to inconsistent optimization objectives. The potential output errors from multiple independently trained models, combined with the signal processing results flowing sequentially between the models in a cascading manner, will lead to error accumulation, ultimately affecting the accuracy of the electronic device's voice interaction response.

[0053] To address the aforementioned issues, the main solution proposed in this application includes: firstly, acquiring audio signals from multiple channels to provide a data foundation for subsequent processing; then, calling a feature acquisition module to acquire a first fusion feature based on the audio signals from multiple channels, achieving preliminary integration of textual and spatial information in the audio signals; next, calling a sound source localization module to acquire a first sound source direction feature based on the first fusion feature, which serves as spatial prior information to enhance spatial perception; then, calling a feature fusion module to fuse the first fusion feature with the first sound source direction feature to obtain a second fusion feature; finally, calling a shared coding module to extract contextual information from the second fusion feature to obtain a first shared feature. The first shared feature includes spatial prior information as well as textual and spatial information from the audio signals, which can subsequently be provided to various speech task modules to execute corresponding speech tasks. Thus, on the one hand, spatial prior information is incorporated into the first shared feature, improving the focusing accuracy and anti-interference capability of target speech in complex environments; on the other hand, end-to-end multi-module collaborative processing overcomes the cumulative errors that may occur in cascaded methods, ultimately effectively improving the accuracy of voice interaction response in electronic devices.

[0054] based on Figure 1 The scene shown below is an illustration; the following will combine... Figures 2-11 This paper provides a detailed description of the signal processing method and model training method provided in the embodiments of this application.

[0055] Please see Figure 2 , Figure 2 This is a schematic flowchart illustrating a signal processing method provided in an embodiment of this application. Figure 2 As shown, the method of this application embodiment is applied to an electronic device. The electronic device has a target model, which includes a feature acquisition module, a sound source localization module, a feature fusion module, a shared coding module, and at least one speech task module. The method of this application embodiment may include the following steps S101-S105.

[0056] S101, acquire audio signals of multiple channels.

[0057] Specifically, the target model involved in the embodiment is a pre-trained end-to-end multi-task neural network model and is deployed in an electronic device. Unlike the cascaded architecture in the related art, the feature acquisition module, the sound source positioning module, the feature fusion module, the shared encoding module, and the at least one speech task module in the target model are a unified whole optimized by joint training, the parameters of each module are shared, and the feature transmission paths are coupled with each other, which can realize multi-task collaborative optimization and end-to-end inference.

[0058] In order to realize voice interaction, first of all, audio signals of multiple channels need to be acquired. The multiple channels refer to signal acquisition channels composed of at least two physically separated microphones; the audio signals of multiple channels refer to original time-domain waveform data or pre-processed frequency-domain feature data containing the same sound source but having phase difference and amplitude difference, which are synchronously acquired by a microphone array.

[0059] It can be understood that the electronic device can acquire multi-channel audio signals through a built-in or external microphone array. The built-in microphone array refers to a microphone array integrated with the main structure of the electronic device and fixedly installed on the device shell; the external microphone array refers to a microphone array connected with the electronic device through a wired or wireless interface and independently arranged from the main body of the device. In addition to the microphone array, the electronic device can also receive pre-recorded multi-channel audio data from a network transmission or a storage device as an input signal source.

[0060] S102, call the feature acquisition module to acquire first fusion features according to the audio signals of multiple channels, and the first fusion features are used to represent the text information and spatial information of the audio signals of multiple channels.

[0061] Specifically, the feature acquisition module refers to a neural network component in the target model responsible for extracting and fusing basic acoustic features from original multi-channel audio signals, and its structure can contain at least one of convolutional layers, recurrent layers, or self-attention mechanism layers.

[0062] To achieve the preliminary feature integration of the multi-channel audio signal, a feature acquisition module needs to be invoked to obtain a first fusion feature according to the audio signals of the multiple channels. The first fusion feature is used to represent the text information and spatial information of the audio signals of the multiple channels. The text information of the audio signals of the multiple channels refers to the semantic information of the speech that can be converted into language content through acoustic modeling, including but not limited to mel-frequency cepstral coefficients, filter bank features or time-frequency spectrum features, and the like acoustic representations. The spatial information of the audio signals of the multiple channels refers to the orientation-related information determined by the microphone array geometry and the sound wave propagation characteristics, including but not limited to phase difference, amplitude difference, time delay difference or beamforming weight, and the like spatial perception representations.

[0063] Regarding the process of invoking the feature acquisition module to obtain the first fusion feature according to the audio signals of the multiple channels, in some possible implementation manners, short-time Fourier transform can be performed on each channel audio signal to obtain a time-frequency spectrum, and then a convolutional neural network is used to extract local acoustic features of each channel, and finally the features of each channel are spliced along the feature dimension to form the first fusion feature. In some possible implementation manners, a generalized cross-correlation function of the multi-channel audio signal can be calculated to obtain a time delay estimation, and then a spatial convolution network is used to extract a directional feature, and finally the directional feature is weighted and fused with the spectral features of each channel to form the first fusion feature containing text information and spatial information.

[0064] The above implementation manners are only examples of "invoking the feature acquisition module to obtain the first fusion feature according to the audio signals of the multiple channels, and the first fusion feature is used to represent the text information and spatial information of the audio signals of the multiple channels", and do not constitute a limitation on other implementation manners.

[0065] In S103, a sound source positioning module is invoked to obtain a first sound source direction feature according to the first fusion feature.

[0066] Specifically, the sound source positioning module refers to a neural network component in the target model that is specifically used to estimate the spatial orientation of the sound source. It distinguishes the direction of the sound source by analyzing the spatial information contained in the first fusion feature. This module can include at least one of a fully connected classification layer, a convolutional neural network or a recurrent neural network.

[0067] To obtain spatial prior information representing the orientation of the sound source to enhance the target focusing ability of the subsequent speech task module, the sound source positioning module needs to be invoked to obtain a first sound source direction feature according to the first fusion feature. The first sound source direction feature refers to a feature representation representing the spatial orientation of the sound source obtained by processing the first fusion feature through the sound source positioning module, which can be in the form of a fixed-dimensional vector. This vector encodes the orientation information of the sound source relative to the microphone array, including but not limited to the feature embedding of spatial parameters such as azimuth angle, elevation angle or distance.

[0068] Regarding the calling of the sound source positioning module, according to the process of obtaining the first sound source direction feature from the first fusion feature, in some possible implementation manners, the first fusion feature can be input into a fully connected layer containing a softmax activation function, and a probability distribution of the sound source direction category is output, and then the probability distribution is mapped into a high-dimensional vector representation through an embedding layer, and finally the vector is taken as the first sound source direction feature. In some possible implementation manners, the first fusion feature can be subjected to spatial feature extraction through a convolutional neural network to generate a spatial feature map, and then the spatial feature map is compressed into a fixed-length feature vector through a global pooling layer, and finally the feature vector is projected into a sound source direction feature space through a linear transformation layer to form the first sound source direction feature. In other possible implementation manners, an attention mechanism can be used to weight and aggregate the spatial information in the first fusion feature to generate an attention weight vector, and then the attention weight vector is multiplied with a learnable sound source direction encoding to finally obtain the first sound source direction feature.

[0069] The above implementation manners only serve as examples of “calling the sound source positioning module and obtaining the first sound source direction feature from the first fusion feature”, and do not constitute a limitation on other implementation manners.

[0070] In S104, a feature fusion module is called to fuse the first fusion feature and the first sound source direction feature to obtain a second fusion feature.

[0071] Specifically, the feature fusion module refers to a neural network component in the target model responsible for integrating features of different sources or different attributes, which realizes deep fusion of information through specific feature combination operations. The feature fusion module can include at least one of a feature concatenation layer, a weighted fusion layer, or an attention mechanism-based fusion layer.

[0072] In order to effectively combine the sound source direction information with the multi-channel audio features to enhance the spatial perception ability of subsequent processing, a feature fusion module is called to fuse the first fusion feature and the first sound source direction feature to obtain a second fusion feature. The second fusion feature refers to a composite feature representation containing text information, spatial information, and sound source direction information of the multi-channel audio signal obtained by combining the first fusion feature and the first sound source direction feature through the feature fusion module.

[0073] The above “fusion” operation refers to a processing process of combining two or more features in the feature dimension, including but not limited to feature concatenation, element-wise addition, element-wise multiplication, or attention mechanism-based weighted fusion.

[0074] Regarding the process of fusing the first fusion feature with the first sound source direction feature to obtain the second fusion feature, in some possible implementation manners, the first sound source direction feature can be dimensionally expanded to match the dimension of the first fusion feature, and then the two features can be connected in the feature dimension through a feature splicing operation to form a composite feature with increased dimension as the second fusion feature. In some possible implementation manners, the first sound source direction feature can be linearly transformed through a learnable weight matrix, and then the transformed feature and the first fusion feature can be added element by element to generate the second fusion feature through a residual connection manner. In another possible implementation manner, an attention mechanism can be adopted, the first sound source direction feature is taken as a query vector, and the first fusion feature is taken as a key-value pair. The second fusion feature is finally obtained by calculating attention weights and performing weighted summation on the first fusion feature, and then fusing the weighted feature with the first sound source direction feature.

[0075] In S105, a shared encoding module is invoked to extract context information from the second fusion feature to obtain a first shared feature, and the first shared feature is used to provide each speech task module in the at least one speech task module to perform a corresponding speech task.

[0076] Specifically, the shared encoding module refers to a core neural network component in the target model responsible for extracting a shared representation with context semantic information from the fusion feature, and the structure of the shared encoding module can include at least one of a recurrent neural network layer, a convolutional neural network layer, a self-attention mechanism layer, or a transformer encoder layer.

[0077] In order to extract a high-level feature representation containing a time sequence dependency and a global context semantic, the shared encoding module needs to be invoked to extract context information from the second fusion feature to obtain the first shared feature. The first shared feature refers to a feature representation containing context information and available for sharing by multiple downstream speech task modules, which is obtained by deep coding the second fusion feature through the shared encoding module.

[0078] The above-mentioned "context information extraction" operation refers to a processing process of capturing long-range dependency relationships and time sequence context associations between features by performing sequence modeling and semantic coding on input features, including but not limited to implementation through hidden state transmission of a recurrent neural network, receptive field expansion of a convolutional neural network, or global relationship modeling of a self-attention mechanism.

[0079] Regarding the calling of the shared coding module, the process of extracting context information from the second fused feature to obtain the first shared feature, in some possible implementation manners, a multi-layer bidirectional long short-term memory network can be used to model the time sequence of the second fused feature, and the first shared feature containing global context information can be obtained by splicing the forward and backward hidden states. In some possible implementation manners, a convolutional neural network can be used to stack multiple convolutional layers to capture context information of different scales by gradually expanding the receptive field, and finally output the first shared feature with multi-level context representation. In other possible implementation manners, a transformer encoder structure can be used to model the global relationship of each time step feature in the second fused feature through a self-attention mechanism, and then transformed through a feedforward neural network to finally generate the first shared feature containing context semantics.

[0080] The speech task module referred to in this embodiment refers to a sub-network component in the target model specially used for performing a specific speech processing function. Exemplarily, the speech task module can be a speech wake-up module, a speech control module, a speech enhancement module, a speech recognition module, a voiceprint recognition module, etc.

[0081] After obtaining the first shared feature, the first shared feature can be further provided as an input of the speech task module, so that the speech task module can perform a corresponding speech processing task based on the shared feature, such as wake-up word detection, instruction word recognition, speech signal enhancement, text transcription, speaker recognition, etc.

[0082] It can be understood that in some cases, the first shared feature can be provided as an input of all speech task modules in the target model to realize multi-task parallel processing and feature sharing. In some cases, the first shared feature can be provided as an input of part of the speech task modules in the target model, and other speech task modules can selectively use or not use the shared feature.

[0083] Finally, the speech task modules process the first shared feature and output corresponding speech task results, and the electronic device performs a corresponding response operation according to the results to realize voice interaction in turn.

[0084] Exemplarily, the shared encoding module can be constructed based on a Conformer (Convolution-augmented Transformer) structure, which can effectively capture global dependency and local context features by organically combining self-attention mechanism and convolution operation. Specifically, the self-attention module in the Conformer structure is responsible for modeling long-range temporal dependencies in the second fusion feature, while the convolution module focuses on extracting detailed information of local acoustic patterns; through the synergistic effect of the feedforward neural network module and the residual connection, multi-scale context information extraction and deep feature transformation of the second fusion feature are realized, and finally the first shared feature with stronger representation ability is output. In addition, the shared encoding module can also be constructed based on a Bidirectional Long Short-Term Memory (BiLSTM) structure, through its unique gating mechanism and bidirectional processing flow, the second fusion feature is modeled from the forward and reverse directions respectively, thereby effectively capturing the forward and backward context dependency in the sequence, and finally the first shared feature containing historical and future context information is output after fusing the hidden states of the two directions.

[0085] In this embodiment, first, the audio signals of multiple channels are acquired to provide a data basis for subsequent processing; then the feature acquisition module is called to acquire the first fusion feature according to the audio signals of multiple channels, realizing the preliminary integration of the text information and spatial information of the audio signals; then the sound source positioning module is called to acquire the first sound source direction feature according to the first fusion feature, and the first sound source direction feature is used as spatial prior information to enhance the spatial perception ability; then the feature fusion module is called to fuse the first fusion feature and the first sound source direction feature to obtain the second fusion feature; finally, the shared encoding module is called to extract the context information of the second fusion feature to obtain the first shared feature, which contains spatial prior information and text information and spatial information of the audio signals, which can be provided to each speech task module to execute the corresponding speech task. In this way, on the one hand, the spatial prior information is integrated into the first shared feature, improving the focusing accuracy and anti-interference ability of the target speech in a complex environment, and on the other hand, through the end-to-end multi-module collaborative processing, the cumulative error that may be generated in the cascading mode is overcome, and finally the accuracy of the electronic device speech interaction response can be effectively improved.

[0086] Please refer to Figure 3 A flowchart for acquiring the first fusion feature is provided for the embodiments of the present application, as shown in Figure 3 The method of the embodiments of the present application can include the following steps S201-S202, which can be further refined as the above-mentioned embodiment step "calling the feature acquisition module to acquire the first fusion feature according to the audio signals of multiple channels".

[0087] S201, calling a feature acquisition module to extract target features corresponding to audio signals of each channel in the audio signals of the plurality of channels, the target features being used to represent text information and spatial information of the audio signals of the corresponding channels;

[0088] S202, fusing the target features corresponding to the audio signals of each channel to obtain first fused features.

[0089] Specifically, the target features refer to high-dimensional feature representations that can simultaneously reflect speech content and spatial attributes, which are extracted from single-channel audio signals by the feature acquisition module. The target features are used to represent text information and spatial information of the audio signals of the corresponding channels, that is, the features not only encode semantic content (text information) such as words and phonemes contained in the speech, but also encode sound wave propagation characteristics (spatial information) determined by the spatial position of the microphone of the channel, such as phase difference, time delay difference, or acoustic cues related to the relative position of the sound source. It can be understood that the target features are one-to-one corresponding to the channels. Assuming that the number of channels is a positive integer N, N target features will also be extracted.

[0090] Regarding the process of calling the feature acquisition module to extract target features corresponding to audio signals of each channel in the audio signals of the plurality of channels, in some possible implementation manners, a short-time Fourier transform can be performed on the audio signals of each channel respectively to convert the time-domain waveform into a time-frequency spectrum, and then a deep processing is performed on the time-frequency spectrum through a set of two-dimensional convolution layers to extract local features containing both frequency-domain acoustic patterns and implicit spatial cues, and a fixed-dimension vector formed after global average pooling of the local features is taken as the target feature corresponding to the audio signals of the channel. In some possible implementation manners, a sequence of mel-frequency cepstral coefficients of each channel audio signal can be calculated first, and then the sequence is input into a one-dimensional convolutional neural network to capture its time-series dynamic change pattern through the sliding of the convolution kernel along the time dimension, and at the same time, the spatial response characteristics related to the channel are implicitly encoded by using the local feature extraction capability of the convolution operation itself, and finally the output features of the one-dimensional convolutional neural network are taken as the target feature corresponding to the audio signals of the channel.

[0091] Further, a feature acquisition module is invoked to fuse the target features corresponding to the audio signals of the channels to obtain a first fused feature. In some possible implementation manners, the target features corresponding to the audio signals of the channels can be spliced along the feature dimension to form a higher-dimensional composite feature vector, and then a fully connected layer is used to reduce the dimension and perform nonlinear transformation on the composite feature vector, and finally the output feature after the transformation is determined as the first fused feature. In some possible implementation manners, an attention fusion mechanism can be used, the target features corresponding to the audio signals of the channels are first input into an attention network, the attention network outputs a weight coefficient corresponding to each channel target feature, and then the target features of the channels are weighted and summed according to the weight coefficients, and the weighted fused feature obtained after the summing is determined as the first fused feature.

[0092] In some possible implementation manners, the feature acquisition module includes an extraction unit and a fusion unit. The extraction unit is configured to extract target features corresponding to the audio signals of the channels in the audio signals of the multiple channels; and the fusion unit is configured to fuse the target features corresponding to the audio signals of the channels to obtain a first fused feature.

[0093] In this embodiment, by respectively extracting independent target features containing text and spatial information in the audio signals of the channels and effectively fusing the target features, it is ensured that the first fused feature can comprehensively and fully retain the original information from all the channels, providing a rich and high-quality feature basis for subsequent sound source positioning and multi-task processing, thereby helping to improve the perception and recognition performance of the entire electronic device in a complex acoustic environment.

[0094] See Figure 4 A flowchart for obtaining a first sound source direction feature is provided for the embodiments of the present application, as shown in Figure 4 The method of the embodiments of the present application can include the following steps S301-S303, which can be further detailed for the above-mentioned embodiment step "invoking a sound source positioning module to obtain a first sound source direction feature according to the first fused feature".

[0095] S301, a sound source positioning module is invoked to perform direction classification processing on the first fused feature to obtain a direction classification result;

[0096] S302, a direction embedding vector corresponding to the direction classification result is generated;

[0097] S303, the direction embedding vector is determined as the first sound source direction feature.

[0098] Specifically, in order to convert the spatial information contained in the first fusion feature into a sound source direction representation with clear physical meaning and enable it to be fused with subsequent modules in a structured vector form, it is first necessary to call a sound source positioning module to perform direction classification processing on the first fusion feature to obtain a direction classification result. The direction classification processing refers to a calculation process of mapping the input first fusion feature to a pre-defined direction category space through a neural network model, which realizes the discretized classification of the sound source direction by pattern recognition and discriminant analysis on the spatial cues related to the sound source direction in the first fusion feature. The direction classification result refers to the data output by the sound source positioning module after processing the first fusion feature, which is used to represent the direction category of the sound source and can be in the form of a one-hot encoding vector, a category probability distribution vector or a category index scalar. Regarding this process, in some possible implementation manners, the first fusion feature can be input to a fully connected classification layer in the sound source positioning module, the number of output neurons of the fully connected classification layer is equal to the number of pre-defined direction categories, and the output is converted into the probability distribution of each category through a softmax activation function, and then the category index with the maximum probability value is determined as the final direction classification result. In some possible implementation manners, a convolutional neural network can be used to extract deep features from the first fusion feature to generate a high-dimensional feature map, and then a global pooling layer is used to compress the feature map into a global feature vector, and finally the global feature vector is input to a linear classifier to output the direction classification result.

[0099] Further, a corresponding direction embedding vector is generated according to the direction classification result. The direction embedding vector refers to a vector representation that can represent the direction semantic information obtained by mapping the discrete direction classification result to a continuous vector space through an embedding layer. The vector encodes the high-level semantic information of the direction category into a fixed-dimensional real number vector form, which facilitates subsequent feature fusion and neural network processing. Regarding this process, in some possible implementation manners, the direction classification result can be used as an index to query a pre-trained direction embedding lookup table to directly obtain the corresponding direction embedding vector from the lookup table. The lookup table is a learnable parameter matrix, the number of rows of which is the same as the number of direction categories, and the number of columns of which is the same as the dimension of the direction embedding vector. In some possible implementation manners, if the direction classification result is a category probability distribution vector, the probability distribution vector can be weighted and summed with all row vectors in the direction embedding lookup table, and the weight is the probability value of the corresponding category. The weighted average vector obtained by the summation is determined as the final direction embedding vector, thereby realizing a soft assignment direction embedding generation method.

[0100] Further, the direction embedding vector is determined as the first sound source direction feature.

[0101] In some possible implementations, the sound source positioning module can be constructed based on a Direction of Arrival (DOA) technique. Specifically, the sound source positioning module can include a combination structure of a beamformer and a classifier, where the beamformer generates a series of beam output energies pointing in different directions by spatially filtering the first fused features, and the classifier determines the sound source direction according to the distribution of the beam output energies and outputs a direction classification result; the sound source positioning module can also use a deep learning-based manner to directly learn the mapping relationship between the sound source direction and the spatial features from the first fused features through a convolutional neural network or a recurrent neural network, and output the direction classification result.

[0102] In this embodiment, by converting the first fused features into discrete direction classification results and then mapping them into continuous direction embedding vectors, not only the semantic information of the sound source direction is preserved, but also the sound source direction can be effectively fused with subsequent features in the form of vectors, the utilization ability of the target model for spatial prior information is enhanced, and explicit spatial guidance is provided for subsequent speech task modules.

[0103] Please refer to Figure 5 A process schematic diagram for introducing the sound source direction after wake-up is provided for the embodiments of the present application, as shown in Figure 5 The method of the embodiments of the present application can include the following steps S401-S404, which can be executed after the above-mentioned embodiment step "calling the feature acquisition module to acquire the first fused features according to the audio signals of multiple channels" and before the step "calling the sound source positioning module to acquire the first sound source direction feature according to the first fused features".

[0104] S401, calling the feature fusion module to fuse the first fused features with the preset second sound source direction feature to obtain third fused features;

[0105] S402, calling the shared coding module to extract the context information of the third fused features to obtain second shared features;

[0106] S403, calling the speech wake-up module to analyze the wake-up instruction according to the second shared features to obtain a wake-up instruction analysis result;

[0107] S404, if the electronic device is controlled to enter the wake-up state according to the wake-up instruction analysis result, the step of calling the sound source positioning module to acquire the first sound source direction feature according to the first fused features is executed.

[0108] Specifically, in order to reduce the false wake-up probability of the electronic device under non-target voice interference and achieve low-power operation, the electronic device can support switching between a wake-up state and a sleep state. The wake-up state refers to a working mode in which the electronic device is ready to receive and execute voice control instructions. In the wake-up state, the voice control module, the voice enhancement module, and other voice task modules of the electronic device are in an activated state and can respond to and process voice instructions of a user. The sleep state refers to a working mode in which the electronic device only maintains basic listening functions to save power. In the sleep state, the electronic device only runs the voice wake-up module to detect specific wake-up words, and other voice task modules are in a low-power or off state.

[0109] Considering that if the first sound source direction feature estimated by the sound source positioning module in real time is introduced when the electronic device is in the sleep state, false wake-up may occur due to interference of environmental noise or non-target speakers. The embodiment proposes that a preliminary wake-up judgment is first performed by using a preset second sound source direction feature, and then accurate sound source direction features are introduced after the wake-up is confirmed, so as to reduce power consumption and computational complexity under the premise of ensuring wake-up accuracy.

[0110] First, the feature fusion module needs to be called to fuse the first fusion feature with the preset second sound source direction feature to obtain a third fusion feature. The second sound source direction feature is a preset feature vector used to replace the sound source direction estimated in real time in the wake-up stage. Its role is to provide a general and non-specific direction spatial priori information for the shared encoding module, so as to realize preliminary spatial perception and wake-up word detection of the multi-channel audio signal without starting the sound source positioning module. Exemplarily, the setting mode of the second sound source direction feature can be a zero vector with all element values of zero, representing no specific direction information; or a placeholder vector with all element values of a specific constant (such as -1), representing unknown direction information; or a learnable embedding vector that can represent the average spatial response or the most common sound source direction obtained through pre-training.

[0111] In some possible implementation manners, the vector value of the second sound source direction feature can be pre-stored in the memory of the electronic device, and the vector value is directly read and fused with the first fusion feature when needed. In some possible implementation manners, a fixed mapping function can be used to generate the corresponding second sound source direction feature according to the current state (such as the sleep state).

[0112] Further, a shared encoding module is called to extract context information from the third fusion feature to obtain a second shared feature. In some possible implementation manners, the shared encoding module can adopt the same network structure and parameters as those for processing the second fusion feature to extract context information from the third fusion feature, so as to ensure consistency of the feature processing procedure. In some possible implementation manners, in order to further reduce the computing overhead in the sleep state, the shared encoding module can adopt a lightweight subnetwork or a simplified model structure to process the third fusion feature, so as to quickly obtain the second shared feature.

[0113] Further, a voice wake-up module is called to perform wake-up instruction analysis according to the second shared feature to obtain a wake-up instruction analysis result. The wake-up instruction analysis result refers to analysis data output by the voice wake-up module after processing the second shared feature, and used to determine whether a preset wake-up word is detected. The data can be expressed as a probability score of the wake-up word, a binary classification result (yes / no wake-up), or a wake-up event marker containing a timestamp. Regarding the process, in some possible implementation manners, the voice wake-up module can include a fully connected classification layer and a softmax activation function, map the second shared feature to a preset wake-up word category space, output a probability of each wake-up word category, and compare the maximum probability with a preset threshold. If the maximum probability exceeds the threshold, it is determined that the corresponding wake-up word is detected. In some possible implementation manners, the voice wake-up module can use a recurrent neural network or a time convolution network to perform sequence modeling on the second shared feature, output a wake-up word detection state at each time step, and then generate a final wake-up instruction analysis result through a post-processing algorithm (such as a sliding window average or a dynamic threshold adjustment).

[0114] It can be understood that the wake-up instruction analysis result can be expressed as detection of a preset wake-up word and an indication that the electronic device should enter a wake-up state, or as non-detection of any preset wake-up word and an indication that the electronic device should maintain the current sleep state.

[0115] If the wake-up instruction analysis result is expressed as detection of a preset wake-up word and an indication that the electronic device should enter a wake-up state, the electronic device can be controlled to enter the wake-up state according to the wake-up instruction analysis result. Specifically, the processor of the electronic device can send a control signal to each voice task module to activate the voice control module, the voice enhancement module, and other modules originally in the sleep state, so that they are ready to receive and process subsequent voice input. Meanwhile, the sampling rate of the microphone array can be increased or additional signal processing functions can be turned on to improve the quality of subsequent voice interaction.

[0116] After the electronic device enters the wake-up state, the step of calling the sound source positioning module according to the first fusion feature to obtain the first sound source direction feature can be performed, so that the accurate sound source direction information is used to enhance the target focusing ability and anti-interference ability of subsequent voice control or voice enhancement tasks.

[0117] In this embodiment, by using the preset second sound source direction feature to replace real-time sound source positioning in the wake-up stage, low-power consumption and low false wake-up rate of wake-up detection are achieved; after successful wake-up, the first sound source direction feature with large calculation amount is enabled, spatial prior information is provided for subsequent accurate voice task processing, and the balance between power consumption and performance is achieved as a whole.

[0118] In an embodiment, the at least one voice task module includes a voice control module; after the above-mentioned embodiment step of "calling the shared coding module to perform context information extraction on the second fusion feature to obtain the first shared feature", the following steps are further included:

[0119] The voice control module is called to perform control instruction analysis according to the first shared feature, and a control instruction analysis result is obtained.

[0120] If a voice control instruction is generated according to the control instruction analysis result, the electronic device is controlled to perform a corresponding voice control function according to the voice control instruction.

[0121] Specifically, to achieve accurate understanding and response to user voice instructions, the voice control module is introduced in this embodiment, which refers to a neural network component in the target model specially used to parse executable control instructions from voice signals. It identifies user intent by analyzing semantic information contained in the first shared feature and outputs corresponding instruction discrimination results. This module can include at least one of a fully connected classification layer, a recurrent neural network, or an attention mechanism layer.

[0122] First, the voice control module needs to be called to perform control instruction analysis according to the first shared feature, and a control instruction analysis result is obtained. The control instruction analysis result refers to data output by the voice control module after processing the first shared feature, which is used to represent the recognized control instruction category. This data can be in the form of a probability distribution vector of preset instruction words, a one-hot encoding of instruction categories, or an instruction label containing a confidence score.

[0123] Regarding the process, in some possible implementations, the first shared feature can be input to a fully connected classification layer, the number of output neurons of the classification layer is equal to the number of preset control instruction categories, and the output is converted into a probability distribution of each category by a softmax activation function, and the instruction corresponding to the category with the maximum probability value is determined as the control instruction analysis result. In some possible implementations, a bidirectional long short-term memory network can be used to model the timing of the first shared feature, capture the timing dynamic change pattern of the instruction word pronunciation, and then output the instruction category probability of each time step through a linear classifier, and finally determine the final control instruction analysis result through a connectionist temporal classification decoding algorithm or a maximum pooling operation.

[0124] It can be understood that the control instruction analysis result can be manifested as successfully identifying a certain preset control instruction and generating a corresponding instruction identifier, or as failing to identify any valid control instruction and outputting an empty instruction or an invalid instruction identifier.

[0125] If the control instruction analysis result is manifested as successfully identifying a certain preset control instruction and generating a corresponding instruction identifier, a voice control instruction can be generated according to the control instruction analysis result. The voice control instruction refers to a standardized instruction format that can be recognized and executed by the control logic of the electronic device according to the conversion of the control instruction analysis result, and the instruction can include instruction type, operation object, parameter value and the like.

[0126] After generating the voice control instruction, the electronic device can be controlled to execute a corresponding voice control function according to the voice control instruction. Exemplarily, the voice control function can be adjusting the temperature setting value of an air conditioning device, switching the play channel of a display device, controlling a sweeping robot to start cleaning work, adjusting the brightness or color temperature of a smart lighting device, controlling the opening and closing degree of a smart curtain, and the like.

[0127] In this embodiment, the first shared feature containing rich context and spatial information is analyzed by the voice control module, the control intention in the user voice can be accurately identified, and an executable voice control instruction can be generated, so that accurate voice control of various functions of the electronic device is realized, and the naturalness and efficiency of human-computer interaction are improved.

[0128] In an embodiment, the voice wake-up module and the voice control module involved in the above embodiments can be integrated in an acoustic model (AM) as a component of the target model. The AM refers to a neural network component in the target model responsible for processing acoustic pattern recognition tasks related to voice content, which detects and classifies specific voice units, keywords or instructions by analyzing acoustic clues in the input features. The acoustic model is trained in an end-to-end manner and optimized jointly with other modules in the target model, which can simultaneously process the wake-up word detection and instruction word recognition tasks and output corresponding voice event discrimination results.

[0129] The voice wake-up module and the voice control module are integrated in the AM, which can specifically be manifested as: the AM internally contains a shared feature extraction subnetwork and task-specific output subnetworks; wherein the shared feature extraction subnetwork is used to extract shared acoustic representations related to voice content from input features; the task-specific output subnetworks include a wake-up output branch and a control instruction output branch, the wake-up output branch is used to calculate the probability of the existence of a wake-up word or generate a wake-up event label according to the shared acoustic representation, and the control instruction output branch is used to identify a preset instruction word category or generate an instruction recognition result according to the shared acoustic representation. In the inference process, the AM can selectively activate the wake-up output branch or the control instruction output branch according to the running state of the electronic device or the external control signal, or perform parallel calculation of the two branches and output the corresponding results.

[0130] In this embodiment, by integrating the voice wake-up module and the voice control module in a unified acoustic model, feature sharing and parameter reuse of wake-up and recognition functions are realized, and the model complexity and computational overhead are reduced; at the same time, due to end-to-end joint training, the shared feature extraction subnetwork in the acoustic model can learn a general acoustic representation suitable for both wake-up and instruction recognition, which improves the performance and efficiency of the model in a multi-task scenario.

[0131] In an embodiment, the at least one voice task module includes a voice enhancement module; after the above embodiment step of "calling the shared encoding module to perform context information extraction on the second fusion feature to obtain the first shared feature", the following steps are further included:

[0132] Calling the voice enhancement module to obtain an enhanced voice signal according to the first shared feature.

[0133] Specifically, in order to improve the quality of the speech signal in a complex acoustic environment and support the cloud advanced semantic understanding function, the speech enhancement module is introduced in the embodiment, which refers to the neural network component in the target model specially used for extracting and reconstructing the pure speech of the target speaker from the noisy or reverberant input signal. The speech enhancement module realizes the enhancement processing of the target speech by utilizing the spatial prior and context information contained in the first shared feature. The module can include at least one of a mask estimation network, a spectral mapping network, or a time domain filtering network.

[0134] Firstly, the speech enhancement module needs to be called to obtain the enhanced speech signal according to the first shared feature. The enhanced speech signal refers to the speech signal representation with improved signal-to-noise ratio and intelligibility obtained by processing the first shared feature through the speech enhancement module. The signal can be in the form of time domain waveform, frequency domain spectrogram, or auditory feature.

[0135] Regarding the process, in some possible implementation manners, the speech enhancement module can adopt an encoder-decoder network based on the U-Net structure, take the first shared feature as the input, gradually compress the feature dimension and capture the advanced semantic information through the encoder path, then gradually up-sample and combine the jump connection to restore the detail information through the decoder path, and finally output the spectral mask of the target speech or the enhanced mel spectrogram as the enhanced speech signal. In some possible implementation manners, the speech enhancement module can adopt a complex spectral mapping method, estimate the real part spectrum and the imaginary part spectrum of the enhanced speech through two parallel fully connected layers respectively, combine the two into a complex spectrum, and reconstruct the enhanced time domain waveform signal as the enhanced speech signal through the inverse short-time Fourier transform.

[0136] In the embodiment, after the shared encoding module is called to extract the context information from the second fusion feature to obtain the first shared feature, the speech enhancement module is further called to obtain the enhanced speech signal according to the first shared feature. The first shared feature contains rich acoustic representations extracted through the context information extraction and the spatial prior fusion. The speech enhancement module can perform targeted signal processing, thereby effectively improving the signal-to-noise ratio and intelligibility of the output speech signal, suppressing environmental noise and reverberation interference, providing a higher quality signal input basis for subsequent speech recognition or transmission, and further improving the overall robustness of the electronic device in a complex acoustic environment.

[0137] In an embodiment, after the above embodiment step of “calling the speech enhancement module to obtain the enhanced speech signal according to the first shared feature”, the following steps are further included:

[0138] The enhanced speech signal is sent to a server in communication connection with the electronic device, so that the server interacts with the electronic device according to the enhanced speech signal.

[0139] Specifically, after obtaining the enhanced speech signal, the enhanced speech signal can be sent to a server in communication connection with the electronic device. The server refers to a remote computer system connected through a network to provide computing resources and complex speech processing functions for the electronic device. The server has strong computing power and rich storage resources, and can run complex speech recognition models, natural language understanding models or dialogue management models.

[0140] Correspondingly, the server receives the enhanced speech signal, and calls a speech recognition engine inside to perform text transcription on the enhanced speech signal, and then analyzes the semantic intention of the transcribed text through a natural language understanding module to finally generate corresponding interaction instructions or response content. Further, the server returns the analyzed interaction instructions or response content to the electronic device, and the electronic device performs corresponding operations or plays corresponding speech feedback according to the returned results to realize the voice interaction function.

[0141] It can be understood that if the electronic device is a household appliance or other type of terminal device used by the user in daily life, the electronic device generally does not have the ability to run large-scale speech recognition models or natural language understanding models. Therefore, more complex operation and analysis tasks can be executed by the server to realize the collaborative working mode of local lightweight processing and cloud high-performance computing, thereby ensuring the real-time response and improving the intelligent level of the interaction experience.

[0142] Based on the communication connection between the electronic device and the server, the embodiment scheme can also realize a hybrid running mode of "offline + online". Specifically, in the offline mode, the electronic device only relies on the local speech wake-up module and the speech control module to process the preset wake-up word and instruction word; in the online mode, the electronic device acquires the enhanced speech signal through the speech enhancement module after being woken up locally and uploads it to the server, and the server completes the complex speech recognition and semantic understanding tasks, thereby realizing flexible switching between offline low power consumption and online high precision.

[0143] In some possible implementations, the enhanced speech signal acquired by the speech enhancement module can also be provided to the speech control module as one of the inputs, that is, the speech control module will analyze the control instruction according to the enhanced speech signal and the first shared feature. In this way, the instruction recognition accuracy and robustness of the speech control module in a strong noise environment can be further improved.

[0144] In some possible implementation manners, the speech enhancement module involved in this embodiment can be constructed by using a speech enhancement (SE) technology. Specifically, the speech enhancement module can be implemented based on a deep learning framework, and a mapping relationship from noisy speech features to clean speech features is learned through training data. The loss function can include a scale-invariant signal-to-noise ratio loss at a time-domain waveform level, a mean square error loss at a frequency-domain spectrogram level, or a perceptual mel- spectrogram adversarial loss at a perceptual level.

[0145] In this embodiment, the high-quality enhanced speech signal generated by the speech enhancement module is uploaded to the server, effectively overcoming the limitation of the local computing capability of the electronic device, and realizing high-precision speech interaction in a complex environment. Meanwhile, the hybrid operation mode of "offline + online" takes into account the low-power and high-performance requirements, and improves the flexibility and usability of the electronic device in actual application.

[0146] The above signal processing method related embodiments are described in detail in the following Figure 6 , Figure 6 is an example schematic diagram of model inference provided by the embodiment of the present application.

[0147] As shown in Figure 6 , the electronic device is provided with a target model trained in advance, and the target model includes a feature acquisition module, a sound source positioning module, a feature fusion module, a shared encoding module, and a speech task module. The speech task module includes a speech wake-up module, a speech control module, and a speech enhancement module.

[0148] First, the electronic device acquires N-channel audio signals synchronously collected by the built-in or external microphone array of the electronic device, and N is a positive integer. The N-channel audio signals contain text information and spatial information of user speech.

[0149] Then, the electronic device calls the feature acquisition module in the target model to process the N-channel audio signals. Specifically, the feature acquisition module extracts target features corresponding to each channel of the N-channel audio signals, and the target features are used to represent the text information and spatial information of the audio signals of the corresponding channel. Then, the feature acquisition module fuses the N target features corresponding to the N-channel audio signals to obtain a first fused feature. The first fused feature is used to represent the text information and spatial information of the N-channel audio signals.

[0150] After successfully obtaining the first fusion feature, in order to perform preliminary wake-up detection and reduce power consumption, the electronic device performs the following steps: calling a feature fusion module, fusing the first fusion feature with a preset second sound source direction feature to obtain a third fusion feature. The second sound source direction feature is a preset placeholder vector, which is used to provide a non-specific spatial prior before wake-up. Then, calling a shared coding module, extracting context information from the third fusion feature to obtain a second shared feature. Next, calling a voice wake-up module, performing wake-up instruction analysis according to the second shared feature to obtain a wake-up instruction analysis result. The wake-up instruction analysis result is used to determine whether a preset wake-up word is detected.

[0151] If it is determined according to the wake-up instruction analysis result that the preset wake-up word is not detected, the electronic device maintains the current state and continues to listen to the audio input. If it is determined according to the wake-up instruction analysis result that the preset wake-up word is detected, the electronic device controls itself to enter a wake-up state, and then performs a subsequent sound source positioning and accurate recognition process.

[0152] After the electronic device enters the wake-up state, a sound source positioning module is called to obtain a first sound source direction feature according to the first fusion feature. The specific process includes: the sound source positioning module performs direction classification processing on the first fusion feature to obtain a direction classification result; a corresponding direction embedding vector is generated according to the direction classification result; and the direction embedding vector is determined as the first sound source direction feature. The first sound source direction feature represents the spatial orientation information of the current main sound source.

[0153] Subsequently, a feature fusion module is called to fuse the first fusion feature with the first sound source direction feature to obtain a second fusion feature. The second fusion feature contains text information, spatial information of the audio signal of the N channels, and spatial prior information represented by the first sound source direction feature.

[0154] Then, a shared coding module is called to extract context information from the second fusion feature to obtain a first shared feature. The first shared feature is a high-level feature representation rich in context information, which integrates audio content, spatial clues, and sound source direction prior.

[0155] After obtaining the first shared feature, the electronic device calls each voice task module in parallel or selectively for processing:

[0156] Firstly, a voice control module is called to perform control instruction analysis according to the first shared feature to obtain a control instruction analysis result. The control instruction analysis result is used to identify a preset control instruction contained in the user voice. If an effective control instruction is successfully identified according to the control instruction analysis result and a corresponding voice control instruction is generated, the electronic device performs a corresponding voice control function according to the voice control instruction, such as adjusting device parameters or triggering a specific operation.

[0157] Secondly, a voice enhancement module is called to obtain an enhanced voice signal according to the first shared feature. The enhanced voice signal is a voice signal after noise reduction and dereverberation and other enhancement processing, and its quality is improved compared with the original audio signal. The electronic device sends the enhanced voice signal to a server in communication connection therewith, so that the server performs further voice recognition and semantic understanding according to the enhanced voice signal, thereby realizing more complex interaction functions, and returns the interaction result to the electronic device.

[0158] It should be noted that the steps of calling the voice control module and calling the voice enhancement module can be performed simultaneously, or can be executed in sequence according to the running mode or strategy of the electronic device. In addition, the enhanced voice signal obtained by the voice enhancement module can also be provided as auxiliary input to the voice control module in some implementation manners, so as to further improve the instruction recognition accuracy of the voice control module in a strong noise environment.

[0159] Through the above process, the embodiment realizes end-to-end multi-task cooperative processing, and effectively improves the performance and robustness of voice wake-up, instruction recognition and voice enhancement of the electronic device in a complex acoustic environment.

[0160] Please refer to Figure 7 , Figure 7 A flowchart of a model training method provided by an embodiment of the present application is shown. As shown in Figure 7 , the method of the embodiment of the present application can include the following steps S501-S507.

[0161] S501, obtaining an initial model, the initial model including a feature acquisition module, a sound source positioning module, a feature fusion module, a shared encoding module, and at least one voice task module.

[0162] Specifically, the embodiment proposes an end-to-end model architecture, which refers to a neural network structure integrating the feature acquisition module, the sound source positioning module, the feature fusion module, the shared encoding module, and the at least one voice task module as a whole, which supports unified modeling and joint optimization from multi-channel audio signal input to multi-task output. The initial model refers to a model structure whose network parameters are in an untrained or randomly initialized state before the training process starts; the feature acquisition module, the sound source positioning module, the feature fusion module, the shared encoding module, and the at least one voice task module included in the initial model have been introduced in the above embodiments, and will not be repeated here.

[0163] S502, obtaining a training sample, the training sample including a plurality of channels of audio signals, labeled information of the sound source positioning module, and labeled information of each voice task module in the at least one voice task module.

[0164] Specifically, the training sample referred to in this embodiment refers to one or more sets of data instances with annotations used to train the initial model. It can be understood that the number of training samples is at least one, and any one training sample includes a plurality of channel audio signals, annotation information of the sound source positioning module, and annotation information of each speech task module.

[0165] The plurality of channel audio signals have been described in the above embodiments, and will not be described here. The annotation information of the sound source positioning module refers to the real sound source direction data used to supervise the training of the sound source positioning module, which can be in the form of direction category label, azimuth value or direction embedding vector, etc. The annotation information of the speech task module refers to the real task related data used to supervise the training of each speech task module, and the specific form of the data depends on the type of the speech task module. For example, for the speech wake-up module, the annotation information can be in the form of wake-up word label, for the speech control module, the annotation information can be in the form of control instruction label, and for the speech enhancement module, the annotation information can be in the form of pure speech signal. It can be understood that the corresponding annotation information is also different for different speech task modules.

[0166] Regarding the process of obtaining the training sample, in some possible implementation manners, the training sample can be constructed by recording multi-channel speech data in a real or simulated acoustic environment and manually adding corresponding annotations, and the training sample is taken as the preset input. In some possible implementation manners, the existing training sample can be expanded through data enhancement technology, and the training sample is taken as the preset input, for example, adding background noise, simulating reverberation effect or changing the sound source direction, to improve the generalization ability of the model.

[0167] S503, calling the feature acquisition module, acquiring the first fusion feature corresponding to the training sample according to the plurality of channel audio signals in the training sample, the first fusion feature corresponding to the training sample being used to represent the text information and the spatial information of the plurality of channel audio signals.

[0168] Specifically, the feature acquisition module and the process of acquiring the first fusion feature by the feature acquisition module in this embodiment have been described in the above signal processing method related embodiments, and will not be described here.

[0169] It should be noted that the first fusion feature corresponding to the training sample is acquired in this embodiment, that is, there is a first fusion feature corresponding to each training sample as the input feature of the subsequent sound source positioning module and the feature fusion module.

[0170] S504, calling the sound source positioning module, acquiring the first sound source direction feature corresponding to the training sample according to the first fusion feature corresponding to the training sample.

[0171] Specifically, the sound source positioning module and the process of obtaining the first sound source direction feature by the sound source positioning module are described in the above-mentioned signal processing method related embodiments, and details are not described herein.

[0172] It should be noted that the first sound source direction feature corresponding to the training sample is obtained in this embodiment, that is, there is a first sound source direction feature corresponding to each training sample as one of the input features of the subsequent feature fusion module.

[0173] S505, calling the feature fusion module, fusing the first fusion feature corresponding to the training sample and the first sound source direction feature to obtain the second fusion feature corresponding to the training sample.

[0174] Specifically, the feature fusion module and the process of obtaining the second fusion feature by the feature fusion module are described in the above-mentioned signal processing method related embodiments, and details are not described herein.

[0175] It should be noted that the second fusion feature corresponding to the training sample is obtained in this embodiment, that is, there is a second fusion feature corresponding to each training sample as an input feature of the subsequent shared encoding module.

[0176] S506, calling the shared encoding module, extracting the context information from the second fusion feature corresponding to the training sample to obtain the first shared feature corresponding to the training sample.

[0177] Specifically, the shared encoding module and the process of obtaining the first shared feature by the shared encoding module are described in the above-mentioned signal processing method related embodiments, and details are not described herein.

[0178] It should be noted that the first shared feature corresponding to the training sample is obtained in this embodiment, that is, there is a first shared feature corresponding to each training sample as an input feature of the subsequent each speech task module and a basis for loss calculation.

[0179] S507, based on the first shared feature corresponding to the training sample, the labeled information of the sound source positioning module and the labeled information of each speech task module, updating the parameters of the initial model to obtain the target model.

[0180] Specifically, after obtaining the first shared feature corresponding to the training sample, the first shared feature corresponding to the training sample can be used as the input of each speech task module, so that each speech task module can execute the speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain the task execution result of each speech task module under the training sample.

[0181] On one hand, the intermediate information generated in the process of obtaining the first sound source direction feature corresponding to the training sample by the sound source positioning module is combined with the labeled information of the sound source positioning module in the training sample, so that the loss parameter of the sound source positioning module under the training sample can be obtained.

[0182] On the other hand, the task execution result of each speech task module under the training sample is combined with the labeled information of each speech task module in the training sample, so that the loss parameter of each speech task module under the training sample can be obtained.

[0183] Further, the loss parameters of the sound source positioning module under the training sample and the loss parameters of each speech task module under the training sample are summarized, and the initial model is updated in a weighted summation, multi-objective optimization or adaptive weight adjustment manner.

[0184] It can be understood that the above parameter updating process is repeatedly executed for multiple training rounds. Each time, the loss parameter is calculated by forward propagation using one or more training samples, and then the gradient is calculated by the back propagation algorithm and the model parameter is updated, until the preset convergence condition is reached, that is, the initial model is determined as the target model. The preset convergence condition can be that the training loss value is stable below a certain threshold, the performance indicator on the validation set no longer improves or reaches the preset maximum number of training rounds.

[0185] In the embodiment, by using the end-to-end joint training method, the loss parameters of the sound source positioning module and each speech task module are considered uniformly, and the collaborative optimization among multiple tasks is realized. Through the loss calculation and parameter updating mechanism based on the first shared feature corresponding to the training sample, the labeled information of the sound source positioning module and the labeled information of each speech task module, it is ensured that the model can learn accurate spatial perception ability and speech task processing ability at the same time. By repeatedly executing the parameter updating process until the preset convergence condition is reached, it is ensured that the target model obtained by training has stable performance and good generalization ability. This training method effectively solves the problem of inconsistent optimization objectives of each model in the related technical cascade structure, and improves the overall performance and robustness of the target model in a complex acoustic environment.

[0186] Please refer to Figure 8 The embodiment of the present application provides a flowchart for obtaining the first fusion feature corresponding to the training sample, as shown in Figure 8 The method of the embodiment of the present application can include the following steps S601-S602, and steps S601-S602 can be further refined as the above-mentioned embodiment step "calling the feature acquisition module, and obtaining the first fusion feature corresponding to the training sample according to the audio signals of multiple channels in the training sample".

[0187] S601, call the feature acquisition module to extract the target features corresponding to the audio signals of each channel in the training sample, and the target features are used to represent the text information and spatial information of the audio signals of the corresponding channel.

[0188] S602, fuse the target features corresponding to the audio signals of each channel in the training sample to obtain the first fused features corresponding to the training sample.

[0189] Specifically, the feature acquisition module, "calling the feature acquisition module to extract the target features corresponding to the audio signals of each channel", and "calling the feature acquisition module to fuse the target features corresponding to the audio signals of each channel in the training sample to obtain the first fused features corresponding to the training sample" have been introduced in the above signal processing method related embodiments. The specific implementation steps of the present embodiment can be obtained by limiting the multi-channel audio signals to specific training samples, which will not be repeated here.

[0190] It should be noted that the present embodiment extracts the target features corresponding to the audio signals of each channel in the training sample, and the subsequent fusion also obtains the first fused features corresponding to the training sample, that is, each training sample corresponds to generate a first fused feature, which will be used as the input feature of the subsequent sound source positioning module and feature fusion module.

[0191] In the present embodiment, by extracting and fusing the target features corresponding to the audio signals of each channel of each training sample, the correspondence between the first fused features and the training samples in the training process is ensured, accurate and consistent input feature representation is provided for the subsequent modules, and the stability and effectiveness of the model training process are guaranteed.

[0192] Please refer to Figure 9 The present embodiment provides a flowchart for obtaining the first sound source direction feature corresponding to the training sample, as shown in Figure 9 The method of the present embodiment can include steps S701-S703, which can be further detailed as the above embodiment step "calling the sound source positioning module to obtain the first sound source direction feature corresponding to the training sample according to the first fused features corresponding to the training sample".

[0193] S701, call the sound source positioning module to perform direction classification processing on the first fused features corresponding to the training sample to obtain the direction classification result corresponding to the training sample;

[0194] S702, generate the direction embedding vector corresponding to the training sample according to the direction classification result corresponding to the training sample;

[0195] S703, determine the direction embedding vector corresponding to the training sample as the first sound source direction feature corresponding to the training sample.

[0196] Specifically, the sound source localization module, "calling the sound source localization module to perform directional classification processing on the first fusion feature corresponding to the training sample to obtain the directional classification result", "generating the corresponding directional embedding vector according to the directional classification result", and "determining the directional embedding vector as the first sound source directional feature corresponding to the training sample" involved in this embodiment have been introduced in the relevant embodiments of the above signal processing method. By establishing the correspondence between the first fusion feature and the specific training sample, the specific implementation steps of this embodiment can be obtained, which will not be repeated here.

[0197] It should be noted that the embodiment obtains the orientation classification result, the orientation embedding vector, and the first sound source orientation feature corresponding to the training sample. That is, each training sample generates an orientation classification result, an orientation embedding vector, and a first sound source orientation feature. These features will serve as the input to the subsequent feature fusion module and the basis for loss calculation.

[0198] In this embodiment, by generating the corresponding orientation classification result, orientation embedding vector and first sound source orientation feature for each training sample, the correspondence between spatial prior information and training samples during training is ensured, providing accurate spatial guidance features for the subsequent feature fusion module, and providing necessary intermediate results for the supervised training of the sound source localization module.

[0199] Please see Figure 10 This application provides a schematic diagram of a process for training a target model, as illustrated in the embodiments of this application. Figure 10 As shown, the method of this application embodiment may include the following steps S801-S805. Steps S801-S805 can be used as a further refinement of the above embodiment step "updating the parameters of the initial model to obtain the target model based on the first shared features corresponding to the training samples, the annotation information of the sound source localization module and the annotation information of each speech task module".

[0200] S801, Based on the direction classification results corresponding to the training samples and the annotation information of the sound source localization module in the training samples, determine the loss parameters of the sound source localization module under the training samples. The direction classification results corresponding to the training samples are obtained in the process of obtaining the first sound source direction features corresponding to the training samples.

[0201] S802, call each speech task module, execute the speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample, and obtain the task execution result of each speech task module under the training sample;

[0202] S803, determine loss parameters of each speech task module under the training sample according to the task execution results of each speech task module under the training sample and the annotation information of each speech task module in the training sample;

[0203] S804, construct a joint loss parameter corresponding to the training sample according to the loss parameter of the sound source positioning module under the training sample and the loss parameters of each speech task module under the training sample;

[0204] S805, update the initial model according to the joint loss parameter corresponding to the training sample to obtain a target model.

[0205] Specifically, in order to realize multi-task joint optimization and overcome the defects of inconsistent optimization goals, error accumulation and limited overall performance caused by multiple independent models due to separate training in the related art, the embodiment proposes an end-to-end parameter updating mechanism based on a unified loss function.

[0206] First, the loss parameter of the sound source positioning module under the training sample needs to be determined according to the direction classification result corresponding to the training sample and the annotation information of the sound source positioning module in the training sample. As for this process, it can be represented as: inputting the direction classification result (for example, a probability distribution vector containing probability values of each category or a specific category index) corresponding to the training sample and the annotation information (for example, a one-hot encoding vector representing the real sound source direction or an azimuth angle value) of the sound source positioning module in the training sample into a preset loss function, and quantifying the prediction error of the sound source positioning module under the current training sample by calculating the difference between the two. The specific form of the loss function can be a cross-entropy loss function (when the direction classification result is a discrete category), a mean square error loss function (when the annotation information is a continuous angle value) or other loss functions suitable for regression or classification tasks. The loss value calculated is the loss parameter of the sound source positioning module under the training sample, which reflects the deviation between the prediction output of the sound source positioning module for the current training sample and the real annotation in numerical value.

[0207] Secondly, each speech task module is called to perform a speech task corresponding to the first shared feature of the training sample, to obtain a task execution result of each speech task module under the training sample. Regarding this process, it can be expressed as: taking the first shared feature corresponding to the training sample as input, feeding it into each speech task module (e.g., a speech wake-up module, a speech control module, and a speech enhancement module) in the initial model. Each speech task module performs forward propagation calculation on the input first shared feature according to its internal preset network structure and current parameters, and outputs a prediction result under its specific task. For example, the speech wake-up module can output a binary flag or a probability score indicating whether a wake-up word is detected; the speech control module can output a probability distribution vector indicating the recognized instruction category; and the speech enhancement module can output an estimated enhanced speech spectrum or time-domain waveform. These outputs generated by each speech task module are the task execution results of each speech task module under the training sample.

[0208] Thirdly, according to the task execution results of each speech task module under the training sample and the labeled information of each speech task module in the training sample, a loss parameter of each speech task module under the training sample is determined. Regarding this process, it can be expressed as: for each speech task module, its task execution result under the training sample and the labeled information of the corresponding speech task module in the training sample (e.g., for the speech wake-up module, its labeled information can be a binary label indicating whether there is a wake-up word; for the speech control module, its labeled information can be a one-hot encoding of an instruction category; and for the speech enhancement module, its labeled information can be a pure speech signal or its feature representation) are input into a loss function preset for the speech task module and matched with the task type. By calculating the difference between the task execution result and the labeled information, the loss value of the speech task module under the current training sample is obtained, i.e., its loss parameter. For example, a cross-entropy loss is commonly used for classification tasks, a mean square error loss is commonly used for regression tasks, and a scale-invariant signal-to-noise ratio loss or a spectral distance loss can be used for speech enhancement tasks.

[0209] Further, according to the loss parameter of the sound source positioning module under the training sample and the loss parameters of each speech task module under the training sample, a joint loss parameter corresponding to the training sample is constructed. Regarding this process, it can be expressed as: the loss parameter of the sound source positioning module under the training sample and the loss parameters of all speech task modules under the training sample are weighted and summed, and the result of the summation is the joint loss parameter corresponding to the training sample. Each loss parameter can be multiplied by a preset weight coefficient before summation, and the weight coefficient is used to balance the contribution of different tasks to the overall optimization goal, and its value can be set according to prior knowledge or determined through hyperparameter optimization.

[0210] Exemplarily, the calculation formula of the weighted sum can be represented as:

[0211] The joint loss parameter = λ1*L_doa + λ2*L_wake + λ3*L_cmd + λ4*L_se

[0212] Wherein, L_doa represents the loss parameter of the sound source positioning module, L_wake, L_cmd, L_se respectively represent the loss parameters of the speech wake-up module, the speech control module, and the speech enhancement module, and λ1, λ2, λ3, λ4 are corresponding weight coefficients.

[0213] It can be understood that since the joint loss parameter integrates the prediction errors of the sound source positioning module and all speech task modules on the same training sample, the joint loss parameter corresponding to the training sample can represent the pros and cons of the overall performance of the initial model on the training sample, which is a scalar value, and the smaller the value is, the better the comprehensive prediction performance of the model on the sample is.

[0214] Finally, according to the joint loss parameter corresponding to the training sample, the parameters of the initial model are updated to obtain the target model. Regarding this process, it can be represented as: using the back propagation algorithm, the gradient of the joint loss parameter with respect to all trainable parameters in the initial model (including the parameters in the feature acquisition module, the sound source positioning module, the feature fusion module, the shared encoding module, and each speech task module) is calculated. Then, using the gradient descent optimization algorithm (or its variants, such as Adam, SGD with momentum, etc.) according to the calculated gradient, the parameters of the initial model are iteratively updated. This process is usually repeated on multiple training samples (a batch), that is, the sum (or average) of the joint loss parameters of all training samples in a batch is calculated, and then the gradient calculation and parameter update are performed according to the total loss. Through repeated iteration of the above process, until the model performance converges on the validation set or reaches the preset training stopping condition (such as the maximum number of iterations), the model obtained at this time is the trained target model.

[0215] In this embodiment, by constructing the joint loss parameter that integrates the multi-task loss and performing unified parameter update based thereon, it is ensured that the feature acquisition module, the sound source positioning module, the feature fusion module, the shared encoding module, and each speech task module can be optimized cooperatively in the training process, and the shared feature representation can meet the needs of sound source positioning and each speech task at the same time, thereby effectively improving the overall performance and robustness of the target model in complex acoustic environments for speech processing.

[0216] In an embodiment, the at least one voice task module comprises a voice wake-up module; the step of "calling each voice task module, and performing a voice task corresponding to each voice task module according to the first shared feature corresponding to the training sample, to obtain a task execution result of each voice task module under the training sample" in the above embodiment can comprise the following steps:

[0217] calling the voice wake-up module, and performing wake-up instruction analysis according to the first shared feature corresponding to the training sample, to obtain a wake-up instruction analysis result of the voice wake-up module under the training sample;

[0218] determining the wake-up instruction analysis result of the voice wake-up module under the training sample as the task execution result of the voice wake-up module under the training sample.

[0219] Specifically, the voice wake-up module, "calling the voice wake-up module, and performing wake-up instruction analysis according to the first shared feature to obtain a wake-up instruction analysis result" in the above embodiment have been described in the above signal processing method related embodiments, and the corresponding relationship between the wake-up instruction analysis result and the specific training sample is established, so that the specific implementation steps of the present embodiment are obtained, which will not be described here.

[0220] Further, the wake-up instruction analysis result of the voice wake-up module under the training sample is determined as the task execution result of the voice wake-up module under the training sample.

[0221] It should be noted that the present embodiment obtains the wake-up instruction analysis result of the voice wake-up module under the training sample, that is, each training sample corresponds to a generated wake-up instruction analysis result, which is used as the specific output of the voice wake-up module under the current training sample for subsequent loss parameter calculation and model optimization. The wake-up instruction analysis result specifically represents the discrimination conclusion for whether the training sample contains the preset wake-up word, which can be a binary classification label, a probability value of the existence of the wake-up word, or a wake-up event marker sequence containing timestamp information, etc.

[0222] Further, after determining the task execution result of the voice wake-up module under the training sample, it needs to be compared with the annotation information of the voice wake-up module in the training sample. The annotation information of the voice wake-up module is the real wake-up state identifier corresponding to the training sample, such as a binary label indicating whether there is a preset wake-up word, or a one-hot encoding vector specifying a wake-up word category. By inputting the task execution result and the annotation information into a preset loss function (such as a binary cross-entropy loss function or a multi-class cross-entropy loss function), the prediction error of the voice wake-up module under the current training sample can be quantified, and then the loss parameter of the module under the training sample is obtained.

[0223] In this embodiment, by generating the corresponding wake-up instruction analysis result for each training sample and explicitly taking it as the task execution result of the voice wake-up module, it is ensured that the supervision signal can accurately guide the parameter optimization direction of the voice wake-up module. This process ensures that the model can effectively learn the acoustic patterns and context features of the wake-up word from the labeled data, improving its detection accuracy and robustness of the wake-up word in the inference stage, while providing a reliable loss parameter source for subsequent construction of multi-task joint loss.

[0224] In an embodiment, the at least one voice task module includes a voice control module; the step of "calling each voice task module, and performing the voice task corresponding to each voice task module according to the first shared feature corresponding to the training sample, to obtain the task execution result of each voice task module under the training sample" in the above embodiment can include the following steps:

[0225] calling the voice control module, and performing control instruction analysis according to the first shared feature corresponding to the training sample, to obtain the control instruction analysis result of the voice control module under the training sample;

[0226] determining the control instruction analysis result of the voice control module under the training sample as the task execution result of the voice control module under the training sample.

[0227] Specifically, the voice control module, "calling the voice control module, and performing control instruction analysis according to the first shared feature to obtain the control instruction analysis result" in the above signal processing method related embodiments have been introduced, and the corresponding relationship between the control instruction analysis result and the specific training sample is established, so that the specific implementation steps of this embodiment are obtained, which will not be repeated here.

[0228] Further, the control instruction analysis result of the voice control module under the training sample is determined as the task execution result of the voice control module under the training sample.

[0229] It should be noted that the control instruction analysis result of the voice control module under the training sample is obtained in this embodiment, that is, a control instruction analysis result is generated for each training sample, which is taken as the specific output representation of the voice control module under the current training sample, and is used for subsequent loss calculation and model parameter optimization. The control instruction analysis result specifically represents the recognition conclusion of the voice instruction contained in the training sample, which can be in the form of a probability distribution vector of a preset instruction word, a one-hot encoding label of an instruction category, or an instruction recognition result sequence containing a confidence score, etc.

[0230] Further, after determining the task execution result of the voice control module under the training sample, it needs to be compared and verified with the annotation information of the voice control module in the training sample. The annotation information of the voice control module is the true control instruction identifier corresponding to the training sample, such as a specified instruction word category label or a semantic action code representing the user's intention. By inputting the task execution result and the annotation information into a preset loss function (such as a cross-entropy loss function commonly used in classification tasks or a connectionist temporal classification loss function), the prediction bias of the voice control module under the current training sample can be accurately quantified, and the loss parameter value of the module under the training sample can be calculated.

[0231] In this embodiment, by generating the corresponding control instruction analysis result for each training sample and explicitly defining it as the task execution result of the voice control module, it is ensured that the supervision signal can accurately guide the parameter update process of the voice control module. This processing flow ensures that the model can effectively learn the acoustic features and semantic patterns of the control instruction from the labeled data, improving its recognition accuracy and robustness for user instructions in the subsequent inference phase, while providing accurate and reliable voice control module loss parameter input for constructing a multi-task joint loss function.

[0232] In an embodiment, the at least one voice task module includes a voice enhancement module; the step of "calling each voice task module, and executing the corresponding voice task of each voice task module according to the first shared feature corresponding to the training sample, to obtain the task execution result of each voice task module under the training sample" in the above embodiment can include the following steps:

[0233] calling the voice enhancement module, and obtaining the enhanced voice signal corresponding to the training sample according to the first shared feature corresponding to the training sample;

[0234] determining the enhanced voice signal corresponding to the training sample as the task execution result of the voice enhancement module under the training sample.

[0235] Specifically, the voice enhancement module and the step of "calling the voice enhancement module and obtaining the enhanced voice signal according to the first shared feature" in the above signal processing method related embodiments have been introduced, and the corresponding relationship between the enhanced voice signal and the specific training sample is established, so that the specific implementation steps of this embodiment can be obtained, which will not be repeated here.

[0236] Further, the enhanced voice signal of the voice enhancement module under the training sample is determined as the task execution result of the voice enhancement module under the training sample.

[0237] It should be noted that the embodiment obtains the enhanced speech signal of the speech enhancement module under the training sample, that is, each training sample corresponds to generate an enhanced speech signal, which is the specific output product of the speech enhancement module under the current training sample, and is used for subsequent loss calculation and model parameter optimization. The enhanced speech signal specifically represents the output result of the original multi-channel audio signal in the training sample after enhancement processing, which can be time domain waveform data after noise reduction and dereverberation processing, enhanced frequency domain spectrum feature, or speech spectrum reconstructed based on mask estimation.

[0238] Further, after determining the task execution result of the speech enhancement module under the training sample, it needs to be compared and analyzed with the labeled information of the speech enhancement module in the training sample. The labeled information of the speech enhancement module is the ideal enhancement target corresponding to the training sample, for example, the original speech signal of the target speaker collected in a pure recording environment, or the standardized pure speech feature processed by artificial processing. By inputting the task execution result and the labeled information into the preset loss function (such as time domain scale invariant signal-to-noise ratio loss, frequency domain mean square error loss, or spectral convergence loss), the signal reconstruction quality of the speech enhancement module under the current training sample can be accurately quantified, and then the loss parameter value of the module under the training sample can be calculated.

[0239] In the embodiment, by generating the corresponding enhanced speech signal for each training sample and explicitly defining it as the task execution result of the speech enhancement module, it is ensured that the supervision signal can accurately guide the parameter update process of the speech enhancement module. The processing flow ensures that the model can effectively learn the acoustic feature processing and signal reconstruction ability required for speech enhancement from the labeled data, improves the enhancement effect and robustness of the model on noisy speech in the subsequent inference stage, and provides accurate and reliable speech enhancement module loss parameter input for constructing a multi-task joint loss function.

[0240] The above model training method related embodiments are summarized as follows: Figure 11 , Figure 11 is an example of a model training provided by the embodiment of the present application.

[0241] As shown in Figure 11 , the embodiment is used to train an end-to-end target model, which will be finally deployed in an electronic device for processing N-channel audio signals collected by a microphone array, N being a positive integer. The target model includes a feature acquisition module, a sound source positioning module, a feature fusion module, a shared encoding module, and at least one speech task module, the speech task module including a speech wake-up module, a speech control module, and a speech enhancement module. The training process of the target model is as follows:

[0242] First, an initial model is obtained. The initial model is a model structure whose network parameters are in an untrained or randomly initialized state, which includes a feature acquisition module, a sound source positioning module, a feature fusion module, a shared encoding module, a voice wake-up module, a voice control module, and a voice enhancement module.

[0243] Subsequently, a training sample set is obtained. Each training sample in the training sample set includes: N-channel audio signals, annotation information of the sound source positioning module, annotation information of the voice wake-up module, annotation information of the voice control module, and annotation information of the voice enhancement module. The N-channel audio signals are multi-channel data containing target speaker speech collected in a known acoustic environment. The annotation information of the sound source positioning module is a real sound source direction label corresponding to the N-channel audio signals, such as an azimuth category. The annotation information of the voice wake-up module is used to indicate whether the preset wake-up word is contained in the audio. The annotation information of the voice control module is used to indicate the specific control instruction category contained in the audio. The annotation information of the voice enhancement module is the target speaker clean speech signal or its feature representation corresponding to the audio collected in a clean environment.

[0244] Next, for each training sample in the training sample set, the following forward propagation and loss calculation process is performed:

[0245] The feature acquisition module is called to obtain the first fusion feature corresponding to the training sample according to the N-channel audio signals in the training sample. The specific process is as follows: the feature acquisition module extracts the target feature corresponding to the audio signal of each channel in the training sample, and the target feature is used to represent the text information and spatial information of the corresponding channel audio signal; then, the N target features extracted are fused to obtain the first fusion feature corresponding to the training sample. The first fusion feature corresponding to the training sample is used to represent the text information and spatial information of the N-channel audio signals in the training sample.

[0246] The sound source positioning module is called to obtain the first sound source direction feature corresponding to the training sample according to the first fusion feature corresponding to the training sample. The specific process is as follows: the sound source positioning module performs direction classification processing on the first fusion feature corresponding to the training sample to obtain the direction classification result corresponding to the training sample; a direction embedding vector corresponding to the training sample is generated according to the direction classification result corresponding to the training sample; and the direction embedding vector corresponding to the training sample is determined as the first sound source direction feature corresponding to the training sample.

[0247] The feature fusion module is called to fuse the first fusion feature corresponding to the training sample and the first sound source direction feature corresponding to the training sample to obtain the second fusion feature corresponding to the training sample.

[0248] Call the shared coding module to perform context information extraction on the second fusion feature corresponding to the training sample, and obtain the first shared feature corresponding to the training sample.

[0249] After obtaining the first shared feature corresponding to the training sample, the following steps are performed in parallel to calculate the loss of each module:

[0250] First, according to the direction classification result corresponding to the training sample (this result has been obtained in the process of obtaining the first sound source direction feature corresponding to the training sample) and the annotation information of the sound source positioning module in the training sample, the loss parameter of the sound source positioning module under the training sample is calculated. For example, the cross-entropy loss function is used to measure the difference between the direction classification result and the true direction label.

[0251] Second, call the voice wake-up module, perform wake-up instruction analysis on the first shared feature corresponding to the training sample, and obtain the wake-up instruction analysis result of the voice wake-up module under the training sample; and determine this wake-up instruction analysis result as the task execution result of the voice wake-up module under the training sample. Then, according to the task execution result and the annotation information of the voice wake-up module in the training sample, the loss parameter of the voice wake-up module under the training sample is calculated, for example, using the binary cross-entropy loss function.

[0252] Third, call the voice control module, perform control instruction analysis on the first shared feature corresponding to the training sample, and obtain the control instruction analysis result of the voice control module under the training sample; and determine this control instruction analysis result as the task execution result of the voice control module under the training sample. Then, according to the task execution result and the annotation information of the voice control module in the training sample, the loss parameter of the voice control module under the training sample is calculated, for example, using the cross-entropy loss function.

[0253] Fourth, call the voice enhancement module, obtain the enhanced speech signal corresponding to the training sample according to the first shared feature corresponding to the training sample; and determine this enhanced speech signal as the task execution result of the voice enhancement module under the training sample. Then, according to the task execution result (enhanced speech signal) and the annotation information (pure speech signal) of the voice enhancement module in the training sample, the loss parameter of the voice enhancement module under the training sample is calculated, for example, using the scale-invariant signal-to-noise ratio loss function or the mean square error loss function.

[0254] Subsequently, according to the loss parameter of the sound source positioning module under the training sample, the loss parameter of the voice wake-up module under the training sample, the loss parameter of the voice control module under the training sample, and the loss parameter of the voice enhancement module under the training sample, the joint loss parameter corresponding to the training sample is constructed. The joint loss parameter is obtained by weighted sum of the above four loss parameters, that is:

[0255] Joint loss parameter = λ1*L_doa + λ2*L_wake + λ3*L_cmd + λ4*L_se

[0256] wherein L_doa, L_wake, L_cmd, L_se are loss parameters of the sound source positioning module, the speech wake-up module, the speech control module, and the speech enhancement module respectively, and λ1, λ2, λ3, λ4 are preset weight coefficients.

[0257] Finally, according to the joint loss parameter corresponding to the training sample, the gradient of all trainable parameters in the initial model is calculated by the back propagation algorithm, and the gradient descent optimization algorithm is used to update the parameters of the initial model.

[0258] The above process is iteratively performed for multiple rounds on the training sample set. In each iteration, the joint loss can be calculated based on one training sample or a batch of training samples, and the parameters are updated until the model performance meets the preset convergence condition (such as the loss on the validation set no longer significantly decreases or reaches the maximum training round). The final model obtained is the trained target model, which can be deployed to the electronic device to perform the above signal processing method.

[0259] Through the above training process, the embodiment realizes the end-to-end joint optimization of the sound source positioning, speech wake-up, speech control, and speech enhancement tasks, ensures that the feature representation shared by each module can meet the needs of different tasks at the same time, and effectively improves the overall performance and generalization ability of the target model in a complex acoustic environment.

[0260] Correspondingly, the embodiment of the application also provides a signal processing apparatus applied to an electronic device, the electronic device being provided with a target model, the target model comprising a feature acquisition module, a sound source positioning module, a feature fusion module, a shared coding module, and at least one speech task module, and the signal processing apparatus comprising an acquisition unit and a signal processing unit.

[0261] The acquisition unit is configured to acquire audio signals of multiple channels.

[0262] The signal processing unit is configured to: invoke the feature acquisition module, acquire first fusion features according to the audio signals of the multiple channels, the first fusion features being used to represent text information and spatial information of the audio signals of the multiple channels; invoke the sound source positioning module, acquire first sound source direction features according to the first fusion features; invoke the feature fusion module, fuse the first fusion features and the first sound source direction features to obtain second fusion features; and invoke the shared coding module, extract context information from the second fusion features to obtain first shared features, the first shared features being used to provide each speech task module in the at least one speech task module to execute a corresponding speech task.

[0263] Optionally, the signal processing unit is further configured to: invoke the feature obtaining module to extract target features corresponding to the audio signals of the channels from the plurality of channels, the target features being used to represent text information and spatial information of the audio signals of the corresponding channels; and fuse the target features corresponding to the audio signals of the channels to obtain first fused features.

[0264] Optionally, the signal processing unit is further configured to: invoke the sound source positioning module to perform directional classification processing on the first fused features to obtain a directional classification result; generate a corresponding directional embedding vector according to the directional classification result; and determine the directional embedding vector as the first sound source directional feature.

[0265] Optionally, the at least one speech task module includes a speech wake-up module, and the signal processing apparatus further includes a wake-up unit; the wake-up unit is configured to: invoke the feature fusion module to fuse the first fused features with a preset second sound source directional feature to obtain third fused features; invoke the shared encoding module to perform context information extraction on the third fused features to obtain second shared features; invoke the speech wake-up module to perform wake-up instruction analysis according to the second shared features to obtain a wake-up instruction analysis result; and if the electronic device is controlled to enter a wake-up state according to the wake-up instruction analysis result, perform the step of invoking the sound source positioning module to obtain the first sound source directional feature according to the first fused features.

[0266] Optionally, the at least one speech task module includes a speech control module, and the signal processing apparatus further includes a speech control unit; the speech control unit is configured to: invoke the speech control module to perform control instruction analysis according to the first shared features to obtain a control instruction analysis result; and if a speech control instruction is generated according to the control instruction analysis result, control the electronic device to perform a corresponding speech control function according to the speech control instruction.

[0267] Optionally, the at least one speech task module includes a speech enhancement module, and the signal processing apparatus further includes a speech enhancement unit; the speech enhancement unit is configured to: invoke the speech enhancement module to obtain an enhanced speech signal according to the first shared features.

[0268] Optionally, the speech enhancement unit is further configured to: send the enhanced speech signal to a server in communication connection with the electronic device, so that the server interacts with the electronic device according to the enhanced speech signal.

[0269] Effects that can be achieved by the embodiments of the present application are described in the above-mentioned related embodiments of the signal processing method, and will not be described here.

[0270] Correspondingly, the embodiments of the present application further provide a model training apparatus, which includes a first obtaining unit, a second obtaining unit, a signal processing unit, and a parameter updating unit.

[0271] The first obtaining unit is configured to obtain an initial model, the initial model comprising a feature obtaining module, a sound source positioning module, a feature fusion module, a shared encoding module, and at least one speech task module.

[0272] The second obtaining unit is configured to obtain a training sample, the training sample comprising audio signals of multiple channels, labeled information of the sound source positioning module, and labeled information of each speech task module in the at least one speech task module.

[0273] The signal processing unit is configured to invoke the feature obtaining module, obtain first fusion features corresponding to the training sample according to the audio signals of the multiple channels in the training sample, the first fusion features corresponding to the training sample being used to represent text information and spatial information of the audio signals of the multiple channels; invoke the sound source positioning module, obtain first sound source direction features corresponding to the training sample according to the first fusion features corresponding to the training sample; invoke the feature fusion module, fuse the first fusion features corresponding to the training sample and the first sound source direction features to obtain second fusion features corresponding to the training sample; and invoke the shared encoding module, extract context information of the second fusion features corresponding to the training sample to obtain first shared features corresponding to the training sample.

[0274] The parameter updating unit is configured to perform parameter updating on the initial model based on the first shared features corresponding to the training sample, the labeled information of the sound source positioning module, and the labeled information of each speech task module to obtain a target model.

[0275] Optionally, the signal processing unit is further configured to invoke the feature obtaining module, extract target features corresponding to the audio signals of each channel in the training sample, the target features being used to represent text information and spatial information of the audio signals of the corresponding channel; and fuse the target features corresponding to the audio signals of each channel in the training sample to obtain the first fusion features corresponding to the training sample.

[0276] Optionally, the signal processing unit is further configured to invoke the sound source positioning module, perform direction classification processing on the first fusion features corresponding to the training sample to obtain direction classification results corresponding to the training sample; generate direction embedding vectors corresponding to the training sample according to the direction classification results corresponding to the training sample; and determine the direction embedding vectors corresponding to the training sample as the first sound source direction features corresponding to the training sample.

[0277] Optionally, the parameter updating unit is further configured to: determine a loss parameter of the sound source positioning module under the training sample according to a direction classification result corresponding to the training sample and the annotation information of the sound source positioning module in the training sample, the direction classification result corresponding to the training sample being obtained in the process of obtaining the first sound source direction feature corresponding to the training sample; invoke each speech task module to execute a speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each speech task module under the training sample; determine a loss parameter of each speech task module under the training sample according to the task execution result of each speech task module under the training sample and the annotation information of each speech task module in the training sample; construct a joint loss parameter corresponding to the training sample according to the loss parameter of the sound source positioning module under the training sample and the loss parameters of each speech task module under the training sample; and perform parameter updating on the initial model according to the joint loss parameter corresponding to the training sample to obtain the target model.

[0278] Optionally, the at least one speech task module includes a speech wake-up module; and the parameter updating unit is further configured to: invoke the speech wake-up module to perform wake-up instruction analysis according to the first shared feature corresponding to the training sample to obtain a wake-up instruction analysis result of the speech wake-up module under the training sample; and determine the wake-up instruction analysis result of the speech wake-up module under the training sample as the task execution result of the speech wake-up module under the training sample.

[0279] Optionally, the at least one speech task module includes a speech control module; and the parameter updating unit is further configured to: invoke the speech control module to perform control instruction analysis according to the first shared feature corresponding to the training sample to obtain a control instruction analysis result of the speech control module under the training sample; and determine the control instruction analysis result of the speech control module under the training sample as the task execution result of the speech control module under the training sample.

[0280] Optionally, the at least one speech task module includes a speech enhancement module; and the parameter updating unit is further configured to: invoke the speech enhancement module to obtain an enhanced speech signal corresponding to the training sample according to the first shared feature corresponding to the training sample; and determine the enhanced speech signal corresponding to the training sample as the task execution result of the speech enhancement module under the training sample.

[0281] Effects that can be achieved by the embodiment are described in the above-mentioned related embodiments of the model training method, which will not be repeated here.

[0282] Correspondingly, the embodiment of the present application further provides an electronic device. Please refer to Figure 12 , Figure 12is a structural schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device 900 includes a processor 901 and a memory 902. The processor 901 is electrically connected to the memory 902.

[0283] The processor 901 is the control center of the electronic device 900, and connects various parts of the entire electronic device 900 through various interfaces and lines, executes various functions of the electronic device 900 and processes data by running or calling executable program codes stored in the memory 902 and calling data stored in the memory 902, thereby monitoring the entire electronic device 900.

[0284] The memory 902 can be used to store executable program codes and modules, and the processor 901 executes signal processing, model training and various functional applications by running the executable program codes and modules stored in the memory 902. The memory 902 can mainly include a program storage area and a data storage area, wherein the program storage area can store executable program codes required by at least one function, etc.; the data storage area can store data created according to the use of the electronic device 900, etc.

[0285] In addition, the memory 902 can include a high-speed random access memory 902, and can also include a non-volatile memory 902, such as at least one magnetic disk memory 902, a flash memory device, or other volatile solid-state memory 902. Accordingly, the memory 902 can also include a memory 902 controller to provide access to the memory 902 by the processor 901.

[0286] In the present embodiment, the processor 901 in the electronic device 900 loads instructions corresponding to the processes of one or more executable program codes into the memory 902, and the processor 901 runs the executable program codes stored in the memory 902, thereby realizing various functions.

[0287] In some possible implementations, the electronic device 900 is provided with a target model, and the target model includes a feature acquisition module, a sound source positioning module, a feature fusion module, a shared coding module, and at least one speech task module. The electronic device 900 can execute a signal processing method, which is specifically executed by the processor 901:

[0288] Obtain audio signals of multiple channels;

[0289] Call the feature acquisition module to obtain first fusion features according to the audio signals of the multiple channels, and the first fusion features are used to represent text information and spatial information of the audio signals of the multiple channels;

[0290] Call the sound source positioning module to obtain first sound source direction features according to the first fusion features;

[0291] The feature fusion module is called to fuse the first fusion feature and the first sound source direction feature to obtain a second fusion feature.

[0292] The shared encoding module is called to extract context information from the second fusion feature to obtain a first shared feature, which is used to provide each speech task module in the at least one speech task module to perform a corresponding speech task.

[0293] Optionally, when the processor 901 executes the calling feature acquisition module to acquire the first fusion feature according to the audio signals of the multiple channels, the processor 901 specifically executes: calling the feature acquisition module to extract target features corresponding to the audio signals of each channel in the audio signals of the multiple channels, the target features being used to represent text information and spatial information of the audio signals of the corresponding channel; and fusing the target features corresponding to the audio signals of each channel to obtain the first fusion feature.

[0294] Optionally, when the processor 901 executes the sound source positioning module to acquire the first sound source direction feature according to the first fusion feature, the processor 901 specifically executes: calling the sound source positioning module to perform direction classification processing on the first fusion feature to obtain a direction classification result; generating a corresponding direction embedding vector according to the direction classification result; and determining the direction embedding vector as the first sound source direction feature.

[0295] Optionally, the at least one speech task module includes a speech wake-up module; after the processor 901 executes the calling feature acquisition module to acquire the first fusion feature according to the audio signals of the multiple channels, before the processor 901 executes the calling sound source positioning module to acquire the first sound source direction feature according to the first fusion feature, the processor 901 can further execute: calling the feature fusion module to fuse the first fusion feature and a preset second sound source direction feature to obtain a third fusion feature; calling the shared encoding module to extract context information from the third fusion feature to obtain a second shared feature; calling the speech wake-up module to analyze a wake-up instruction according to the second shared feature to obtain a wake-up instruction analysis result; and if the electronic device 900 is controlled to enter a wake-up state according to the wake-up instruction analysis result, executing the step of calling the sound source positioning module to acquire the first sound source direction feature according to the first fusion feature.

[0296] Optionally, the at least one speech task module includes a speech control module; after the processor 901 executes the calling shared encoding module to extract context information from the second fusion feature to obtain the first shared feature, the processor 901 can further execute: calling the speech control module to analyze a control instruction according to the first shared feature to obtain a control instruction analysis result; and if a speech control instruction is generated according to the control instruction analysis result, controlling the electronic device 900 to perform a corresponding speech control function according to the speech control instruction.

[0297] Optionally, the at least one speech task module comprises a speech enhancement module; after the processor 901 executes the calling of the shared coding module to perform context information extraction on the second fusion feature to obtain the first shared feature, the processor 901 can further execute: calling the speech enhancement module to obtain an enhanced speech signal according to the first shared feature; and sending the enhanced speech signal to a server in communication connection with the electronic device 900, so that the server interacts with the electronic device 900 according to the enhanced speech signal.

[0298] In some possible implementation manners, the electronic device 900 can perform a model training method, specifically performed by the processor 901:

[0299] obtaining an initial model, the initial model comprising a feature acquisition module, a sound source positioning module, a feature fusion module, a shared coding module, and at least one speech task module;

[0300] obtaining a training sample, the training sample comprising audio signals of multiple channels, labeled information of the sound source positioning module, and labeled information of each speech task module in the at least one speech task module;

[0301] calling the feature acquisition module to acquire, according to the audio signals of the multiple channels in the training sample, a first fusion feature corresponding to the training sample, the first fusion feature corresponding to the training sample being used to represent text information and spatial information of the audio signals of the multiple channels;

[0302] calling the sound source positioning module to acquire, according to the first fusion feature corresponding to the training sample, a first sound source direction feature corresponding to the training sample;

[0303] calling the feature fusion module to fuse the first fusion feature corresponding to the training sample and the first sound source direction feature to obtain a second fusion feature corresponding to the training sample;

[0304] calling the shared coding module to perform context information extraction on the second fusion feature corresponding to the training sample to obtain a first shared feature corresponding to the training sample;

[0305] based on the first shared feature corresponding to the training sample, the labeled information of the sound source positioning module, and the labeled information of each speech task module, performing parameter updating on the initial model to obtain a target model.

[0306] Optionally, when the processor 901 executes the calling of the feature acquisition module to acquire, according to the audio signals of the multiple channels in the training sample, the first fusion feature corresponding to the training sample, the processor 901 specifically executes: calling the feature acquisition module to extract a target feature corresponding to the audio signal of each channel in the training sample, the target feature being used to represent text information and spatial information of the audio signal of the corresponding channel; and fusing the target features corresponding to the audio signals of the multiple channels in the training sample to obtain the first fusion feature corresponding to the training sample.

[0307] Optionally, the processor 901, in the execution of calling the sound source positioning module, and obtaining the first sound source direction feature corresponding to the training sample according to the first fusion feature corresponding to the training sample, specifically executes: calling the sound source positioning module, performing direction classification processing on the first fusion feature corresponding to the training sample to obtain a direction classification result corresponding to the training sample; generating a direction embedding vector corresponding to the training sample according to the direction classification result corresponding to the training sample; and determining the direction embedding vector corresponding to the training sample as the first sound source direction feature corresponding to the training sample.

[0308] Optionally, the processor 901, in the execution of updating the initial model to obtain the target model based on the first shared feature corresponding to the training sample, the labeled information of the sound source positioning module, and the labeled information of each speech task module, specifically executes: determining a loss parameter of the sound source positioning module under the training sample according to the direction classification result corresponding to the training sample and the labeled information of the sound source positioning module in the training sample, the direction classification result corresponding to the training sample being obtained in the process of obtaining the first sound source direction feature corresponding to the training sample; calling each speech task module, and executing a speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each speech task module under the training sample; determining a loss parameter of each speech task module under the training sample according to the task execution result of each speech task module under the training sample and the labeled information of each speech task module in the training sample; constructing a joint loss parameter corresponding to the training sample according to the loss parameter of the sound source positioning module under the training sample and the loss parameters of each speech task module under the training sample; and updating the initial model to obtain the target model according to the joint loss parameter corresponding to the training sample.

[0309] Optionally, the at least one speech task module includes a speech wake-up module; and the processor 901, in the execution of calling each speech task module, executing a speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each speech task module under the training sample, specifically executes: calling the speech wake-up module, performing wake-up instruction analysis according to the first shared feature corresponding to the training sample to obtain a wake-up instruction analysis result of the speech wake-up module under the training sample; and determining the wake-up instruction analysis result of the speech wake-up module under the training sample as the task execution result of the speech wake-up module under the training sample.

[0310] Optionally, the at least one voice task module includes a voice control module; when the processor 901 executes calling each voice task module, performing a voice task corresponding to each voice task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each voice task module under the training sample, the processor 901 specifically executes: calling the voice control module, performing control instruction analysis according to the first shared feature corresponding to the training sample to obtain a control instruction analysis result of the voice control module under the training sample; and determining the control instruction analysis result of the voice control module under the training sample as the task execution result of the voice control module under the training sample.

[0311] Optionally, the at least one voice task module includes a voice enhancement module; when the processor 901 executes calling each voice task module, performing a voice task corresponding to each voice task module according to the first shared feature corresponding to the training sample to obtain a task execution result of each voice task module under the training sample, the processor 901 specifically executes: calling the voice enhancement module, obtaining an enhanced voice signal corresponding to the training sample according to the first shared feature corresponding to the training sample; and determining the enhanced voice signal corresponding to the training sample as the task execution result of the voice enhancement module under the training sample.

[0312] Effects that can be achieved by the embodiments of the present application are described in the above-mentioned signal processing method and model training method, which will not be repeated here.

[0313] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. When the computer program is executed, the computer program causes a computer to execute the above-mentioned related method steps to implement any one of the signal processing methods or model training methods provided by the above-mentioned embodiments.

[0314] The embodiments of the present application further provide a computer program product, which stores at least one instruction. When the at least one instruction is executed by a processor, the at least one instruction implements any one of the signal processing methods or model training methods provided by the above-mentioned embodiments.

[0315] The signal processing device, the model training device, the electronic device, the computer readable storage medium, and the computer program product provided by the embodiments of the present application are used to execute the corresponding methods provided above, and thus the beneficial effects achieved by the signal processing device, the model training device, the electronic device, the computer readable storage medium, and the computer program product can refer to the beneficial effects of the corresponding methods provided above, which will not be repeated here.

[0316] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A signal processing method, characterized by, The method is applied to an electronic device provided with a target model, and the target model comprises a feature acquisition module, a sound source positioning module, a feature fusion module, a shared coding module and at least one speech task module, and the method comprises: obtaining audio signals of multiple channels; calling the feature acquisition module to obtain first fusion features according to the audio signals of the multiple channels, the first fusion features being used to represent text information and spatial information of the audio signals of the multiple channels; calling the sound source positioning module to obtain first sound source direction features according to the first fusion features; calling the feature fusion module to fuse the first fusion features and the first sound source direction features to obtain second fusion features; calling the shared coding module to extract context information from the second fusion features to obtain first shared features, the first shared features being used to provide each of the speech task modules to perform a corresponding speech task.

2. The method of claim 1, wherein, The calling of the feature acquisition module to obtain first fusion features according to the audio signals of the multiple channels comprises: calling the feature acquisition module to extract target features corresponding to the audio signals of each of the multiple channels from the audio signals of the multiple channels, the target features being used to represent text information and spatial information of the audio signals of the corresponding channels; fusing the target features corresponding to the audio signals of each of the multiple channels to obtain the first fusion features.

3. The method of claim 1, wherein, The calling of the sound source positioning module to obtain first sound source direction features according to the first fusion features comprises: calling the sound source positioning module to perform direction classification processing on the first fusion features to obtain a direction classification result; generating a corresponding direction embedding vector according to the direction classification result; determining the direction embedding vector as the first sound source direction features.

4. The method according to any one of claims 1 to 3, characterized in that, The at least one speech task module comprises a speech wake-up module; after the calling of the feature acquisition module to obtain first fusion features according to the audio signals of the multiple channels, before the calling of the sound source positioning module to obtain first sound source direction features according to the first fusion features, the method further comprises: calling the feature fusion module to fuse the first fusion features and a preset second sound source direction feature to obtain third fusion features; calling the shared coding module to extract context information from the third fusion features to obtain second shared features; calling the speech wake-up module to analyze a wake-up instruction according to the second shared features to obtain a wake-up instruction analysis result; if the electronic device is controlled to enter a wake-up state according to the wake-up instruction analysis result, the calling of the sound source positioning module to obtain first sound source direction features according to the first fusion features is performed.

5. The method according to any one of claims 1 to 3, characterized in that, The at least one speech task module comprises a speech control module; after the calling of the shared coding module to extract context information from the second fusion features to obtain first shared features, the method further comprises: calling the speech control module to analyze a control instruction according to the first shared features to obtain a control instruction analysis result; If a voice control instruction is generated according to the control instruction analysis result, the electronic device is controlled to execute a corresponding voice control function according to the voice control instruction.

6. The method according to any one of claims 1 to 3, characterized in that, The at least one voice task module includes a voice enhancement module; after the shared coding module is invoked to extract the context information from the second fusion feature to obtain the first shared feature, the method further includes: The voice enhancement module is invoked to obtain an enhanced voice signal according to the first shared feature.

7. The method of claim 6, wherein, After the voice enhancement module is invoked to obtain an enhanced voice signal according to the first shared feature, the method further includes: The enhanced voice signal is sent to a server in communication connection with the electronic device, so that the server interacts with the electronic device according to the enhanced voice signal.

8. A model training method, comprising: The method includes: An initial model is obtained, the initial model including a feature acquisition module, a sound source positioning module, a feature fusion module, a shared coding module, and at least one voice task module; A training sample is obtained, the training sample including audio signals of multiple channels, labeled information of the sound source positioning module, and labeled information of each of the at least one voice task module; The feature acquisition module is invoked to obtain first fusion features corresponding to the training sample according to the audio signals of multiple channels in the training sample, the first fusion features corresponding to the training sample being used to represent text information and spatial information of the audio signals of multiple channels; The sound source positioning module is invoked to obtain first sound source direction features corresponding to the training sample according to the first fusion features corresponding to the training sample; The feature fusion module is invoked to fuse the first fusion features corresponding to the training sample and the first sound source direction features to obtain second fusion features corresponding to the training sample; The shared coding module is invoked to extract context information from the second fusion features corresponding to the training sample to obtain first shared features corresponding to the training sample; Based on the first shared features corresponding to the training sample, the labeled information of the sound source positioning module, and the labeled information of each of the at least one voice task module, the initial model is parameter updated to obtain a target model.

9. The method of claim 8, wherein, The feature acquisition module is invoked to obtain first fusion features corresponding to the training sample according to the audio signals of multiple channels in the training sample, including: The feature acquisition module is invoked to extract target features corresponding to the audio signals of each channel in the training sample, the target features being used to represent text information and spatial information of the audio signals of the corresponding channel; The target features corresponding to the audio signals of each channel in the training sample are fused to obtain the first fusion features corresponding to the training sample.

10. The method of claim 8, wherein, The sound source positioning module is invoked to obtain first sound source direction features corresponding to the training sample according to the first fusion features corresponding to the training sample, including: The sound source positioning module is invoked to perform direction classification processing on the first fusion features corresponding to the training sample to obtain direction classification results corresponding to the training sample; A direction embedding vector corresponding to the training sample is generated according to the direction classification results corresponding to the training sample; The direction embedding vector corresponding to the training sample is determined as the first sound source direction feature corresponding to the training sample.

11. The method according to any one of claims 8 to 10, characterized in that, The parameter updating of the initial model to obtain a target model based on the first shared feature corresponding to the training sample, the labeled information of the sound source positioning module, and the labeled information of each speech task module, comprises: According to the direction classification result corresponding to the training sample and the labeled information of the sound source positioning module in the training sample, the loss parameter of the sound source positioning module under the training sample is determined, which is obtained in the process of obtaining the first sound source direction feature corresponding to the training sample; Each speech task module is called, and each speech task corresponding to each speech task module is executed according to the first shared feature corresponding to the training sample, to obtain the task execution result of each speech task module under the training sample; According to the task execution result of each speech task module under the training sample and the labeled information of each speech task module in the training sample, the loss parameter of each speech task module under the training sample is determined; According to the loss parameter of the sound source positioning module under the training sample and the loss parameter of each speech task module under the training sample, the joint loss parameter corresponding to the training sample is constructed; According to the joint loss parameter corresponding to the training sample, the parameter updating of the initial model is performed to obtain a target model.

12. The method of claim 11, wherein, At least one of the speech task modules comprises a speech wake-up module; the calling of each speech task module to execute each speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain the task execution result of each speech task module under the training sample, comprises: The speech wake-up module is called to perform wake-up instruction analysis according to the first shared feature corresponding to the training sample, to obtain the wake-up instruction analysis result of the speech wake-up module under the training sample; The wake-up instruction analysis result of the speech wake-up module under the training sample is determined as the task execution result of the speech wake-up module under the training sample.

13. The method of claim 11, wherein, At least one of the speech task modules comprises a speech control module; the calling of each speech task module to execute each speech task corresponding to each speech task module according to the first shared feature corresponding to the training sample to obtain the task execution result of each speech task module under the training sample, comprises: The speech control module is called to perform control instruction analysis according to the first shared feature corresponding to the training sample, to obtain the control instruction analysis result of the speech control module under the training sample; The control instruction analysis result of the speech control module under the training sample is determined as the task execution result of the speech control module under the training sample.

14. The method of claim 11, wherein, The at least one voice task module comprises a voice enhancement module; the voice task module is invoked to perform a voice task corresponding to the voice task module according to the first shared feature corresponding to the training sample, to obtain a task execution result of the voice task module under the training sample, comprising: The voice enhancement module is invoked to obtain an enhanced voice signal corresponding to the training sample according to the first shared feature corresponding to the training sample; The enhanced voice signal corresponding to the training sample is determined as the task execution result of the voice enhancement module under the training sample.

15. A signal processing device, characterized by The signal processing device is applied to an electronic device, and the electronic device is provided with a target model, the target model comprising a feature acquisition module, a sound source positioning module, a feature fusion module, a shared encoding module, and at least one voice task module; the signal processing device comprises an acquisition unit and a signal processing unit; The acquisition unit is configured to acquire audio signals of multiple channels; The signal processing unit is configured to invoke the feature acquisition module to acquire first fusion features according to the audio signals of the multiple channels, the first fusion features being used to represent text information and spatial information of the audio signals of the multiple channels; The sound source positioning module is invoked to acquire first sound source direction features according to the first fusion features; The feature fusion module is invoked to fuse the first fusion features and the first sound source direction features to obtain second fusion features; The shared encoding module is invoked to extract context information from the second fusion features to obtain first shared features, the first shared features being used to provide each voice task module in the at least one voice task module to perform a corresponding voice task.

16. A model training apparatus, comprising: The model training device comprises a first acquisition unit, a second acquisition unit, a signal processing unit, and a parameter updating unit; The first acquisition unit is configured to acquire an initial model, the initial model comprising a feature acquisition module, a sound source positioning module, a feature fusion module, a shared encoding module, and at least one voice task module; The second acquisition unit is configured to acquire a training sample, the training sample comprising audio signals of multiple channels, labeled information of the sound source positioning module, and labeled information of each voice task module in the at least one voice task module; The signal processing unit is configured to invoke the feature acquisition module to acquire first fusion features corresponding to the training sample according to the audio signals of the multiple channels in the training sample, the first fusion features corresponding to the training sample being used to represent text information and spatial information of the audio signals of the multiple channels; The sound source positioning module is invoked to acquire first sound source direction features corresponding to the training sample according to the first fusion features corresponding to the training sample; The feature fusion module is invoked to fuse the first fusion features corresponding to the training sample and the first sound source direction features to obtain second fusion features corresponding to the training sample; The shared coding module is called to perform context information extraction on the second fusion feature corresponding to the training sample, to obtain a first shared feature corresponding to the training sample; The parameter updating unit is configured to perform parameter updating on the initial model based on the first shared feature corresponding to the training sample, the labeled information of the sound source positioning module, and the labeled information of each speech task module, to obtain a target model.

17. An electronic device, comprising: The electronic device includes: a memory for storing executable program code; a processor for calling and running the executable program code from the memory, so that the electronic device executes the method of any one of claims 1 to 14.

18. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which, when executed, implements the method of any one of claims 1 to 14.

19. A computer program product storing at least one instruction, which, when executed by a processor, implements the method of any one of claims 1 to 14.