Sound Localization and Recognition Method and Device Based on Positional Encoding Convolutional Neural Network
The PCNN method addresses the issue of time-position interference in traditional CNNs by employing a multi-task model with encoding and decoding stages to enhance sound event localization and recognition accuracy through feature extraction and Transformer models.
Patent Information
- Application Number
- CN202111654890.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-30
AI Technical Summary
In the location and recognition of sound events, traditional convolutional neural networks have poor retention capabilities for relative position information, resulting in temporal position information interference in feature extraction, affecting the accuracy of positioning and recognition results.
The position-coded convolutional neural network is adopted to encode the position information of the target sound source signal through the encoding model, and feature extraction and decoding are combined with the multi-task model to eliminate interference with the time position information, and to use multi-task learning to improve positioning and recognition accuracy.
Effectively eliminate the interference of time position information in the feature vector, improve the positioning accuracy and recognition accuracy of sound events, and enhance the robustness and adaptability of the positioning and recognition model.
Smart Images

Figure CN114420150B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and particularly to a sound localization and recognition method and device based on a position-encoding convolutional neural network. Background Art
[0002] The tasks of sound event localization and recognition are to effectively localize the sound source and identify the category of the sound source for various sound events that occur continuously or intermittently and randomly in continuous audio signals. In recent years, artificial intelligence technologies represented by deep learning have been widely applied in various fields, and the field of audio signal processing is no exception.
[0003] Currently, as a deep neural network with extremely strong representation ability, the traditional convolutional neural network is integrated into various audio signal processing algorithms. In the sound event localization and recognition algorithm based on the convolutional neural network, the traditional convolutional neural network is widely used in the feature extraction stage of audio signal processing. Due to the traditional convolutional neural network having a certain displacement invariance, the features of various sound events can be extracted relatively effectively.
[0004] Although the prior art uses the traditional convolutional neural network to localize and recognize sound events, showing excellent performance. However, due to the poor ability of the traditional convolutional neural network to retain relative position information, the interference of the time position information of the occurrence of the sound event cannot be effectively eliminated in the features extracted by the traditional convolutional neural network, resulting in the interference of time position information in the features extracted by the traditional convolutional neural network, thus making it difficult to ensure the accuracy of the sound event localization result and the recognition result. Summary of the Invention
[0005] The present invention provides a sound localization and recognition method and device based on a position-encoding convolutional neural network, which is used to solve the defect that the features extracted by the traditional convolutional neural network in the prior art have interference of time position information, resulting in inaccurate sound event localization results and recognition results, and to achieve accurate localization and recognition of sound events.
[0006] The present invention provides a sound localization and recognition method based on a position-encoding convolutional neural network, including:
[0007] Inputting a target sound source signal into an encoding model in a multi-task model to obtain an encoding result of the target sound source signal; wherein, the encoding model is used to perform position information encoding on the target sound source signal;
[0008] Inputting the target sound source signal and the encoding result into a feature extraction model in the multi-task model to obtain a feature vector of the target sound source signal;
[0009] Input the feature vector of the target sound source signal into the decoding model in the multi-task model to obtain the decoding result of the target sound source signal;
[0010] Input the decoding result of the target sound source signal into the localization and recognition model in the multi-task model to obtain the localization result and recognition result of the target sound source signal;
[0011] Among them, the multi-task model is trained based on the sample sound source signal and the corresponding reference localization result and reference recognition result.
[0012] According to a sound source localization and recognition method based on a position-encoding convolutional neural network provided by the present invention, the feature extraction model includes a first sub-feature extraction model and a second sub-feature extraction model, and the feature vector includes a first sub-feature vector and a second sub-feature vector;
[0013] Correspondingly, the step of inputting the target sound source signal and the encoding result into the feature extraction model in the multi-task model to obtain the feature vector of the target sound source signal includes:
[0014] Input the target sound source signal and the encoding result into the first sub-feature extraction model to obtain the first sub-feature vector of the target sound source signal, and input the target sound source signal and the encoding result into the second sub-feature extraction model to obtain the second sub-feature vector of the target sound source signal;
[0015] Among them, the first sub-feature extraction model is used to extract features related to the localization result of the target sound source signal, and the second sub-feature extraction model is used to extract features related to the recognition result of the target sound source signal.
[0016] According to a sound source localization and recognition method based on a position-encoding convolutional neural network provided by the present invention, the feature extraction model includes at least one group of position information preservation modules and pooling modules;
[0017] The position information preservation module includes a plurality of first convolutional modules with different scales and a second convolutional module;
[0018] The plurality of first convolutional modules with different scales are used to perform multi-scale feature extraction on the target sound source signal and the encoding result to obtain a plurality of feature vectors of the target sound source signal with different scales;
[0019] The second convolutional module is used to fuse the feature vectors with different scales;
[0020] The pooling module is used to perform a pooling operation on the fusion result.
[0021] A sound localization and recognition method based on a position-encoding convolutional neural network according to the present invention, wherein the localization and recognition model includes at least one set of parallel first Transformer models and second Transformer models;
[0022] Each set of the first Transformer models is used to localize each sound event of the target sound source signal;
[0023] Each set of the second Transformer models is used to recognize each sound event of the target sound source signal.
[0024] A sound localization and recognition method based on a position-encoding convolutional neural network according to the present invention, before inputting the target sound source signal into the encoding model of the multi-task model to obtain the encoding result of the target sound source signal, further includes:
[0025] After performing preliminary data augmentation on the sample sound source signal, perform preliminary feature extraction to obtain the preliminary feature vector of the sample sound source signal;
[0026] And / or, perform secondary data augmentation on some of the feature vectors in the preliminary feature vector of the sample sound source signal;
[0027] Train the multi-task model according to the preliminary feature vector of the sample sound source signal and / or the partially feature vectors after secondary data augmentation, as well as the reference localization result and reference recognition result corresponding to the sample sound source signal.
[0028] A sound localization and recognition method based on a position-encoding convolutional neural network according to the present invention, the preliminary feature vector includes a logarithmic Mel spectrogram feature vector and an intensity feature vector;
[0029] Correspondingly, performing secondary data augmentation on some of the feature vectors in the preliminary feature vector of the sample sound source signal includes:
[0030] Perform Mel spectrogram data augmentation on the logarithmic Mel spectrogram feature vector in the preliminary feature vector.
[0031] A sound localization and recognition method based on a position-encoding convolutional neural network according to the present invention, the preliminary data augmentation includes rotating the sample sound source signal in one or more directions, and / or performing random superposition data augmentation on different categories of sample sound source signals.
[0032] The present invention also provides a sound localization and recognition device based on a position-encoding convolutional neural network, including:
[0033] An encoding module, configured to input a target sound source signal into an encoding model in a multi-task model to obtain an encoding result of the target sound source signal; wherein, the encoding model is used to perform position information encoding on the target sound source signal;
[0034] A feature extraction module, configured to input the target sound source signal and the encoding result into a feature extraction model in the multi-task model to obtain a feature vector of the target sound source signal;
[0035] A decoding module, configured to input the feature vector of the target sound source signal into a decoding model in the multi-task model to obtain a decoding result of the target sound source signal;
[0036] A positioning and recognition module, configured to input the decoding result of the target sound source signal into a positioning and recognition model in the multi-task model to obtain a positioning result and a recognition result of the target sound source signal;
[0037] Wherein, the multi-task model is trained based on a sample sound source signal and a reference positioning result and a reference recognition result corresponding to the sample sound source signal.
[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the sound source positioning and recognition method based on a position encoding convolutional neural network as described in any one of the above are implemented.
[0039] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the sound source positioning and recognition method based on a position encoding convolutional neural network as described in any one of the above are implemented.
[0040] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the sound source positioning and recognition method based on a position encoding convolutional neural network as described in any one of the above are implemented.
[0041] The sound source positioning and recognition method and device based on a position encoding convolutional neural network provided by the present invention. On the one hand, by performing position information encoding on the target sound source signal and then performing feature extraction, the interference of time position information in the feature vector is eliminated, and the essential features affecting the positioning task and the recognition task are deeply mined from the target sound source signal, thereby effectively improving the positioning accuracy and recognition accuracy of the target sound source signal. On the other hand, a multi-task model is used to jointly learn the positioning task and the recognition task at the same time, fully considering the correlation and difference between the positioning task and the recognition task, and further improving the positioning accuracy and recognition accuracy of the target sound source signal. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0043] Figure 1 It is one of the schematic flowcharts of the sound localization and recognition method based on the position encoding convolutional neural network provided by the present invention;
[0044] Figure 2 It is the schematic structural diagram of the feature extraction model in the sound localization and recognition method based on the position encoding convolutional neural network provided by the present invention;
[0045] Figure 3 It is the schematic structural diagram of the position information retention module in the sound localization and recognition method based on the position encoding convolutional neural network provided by the present invention;
[0046] Figure 4 It is the second schematic flowchart of the sound localization and recognition method based on the position encoding convolutional neural network provided by the present invention;
[0047] Figure 5 It is the schematic structural diagram of the sound localization and recognition device based on the position encoding convolutional neural network provided by the present invention;
[0048] Figure 6 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0049] To make the purpose, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0050] In the description of the present application, the related descriptions such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0051] The following combines Figure 1 to describe the sound localization and recognition method based on the position encoding convolutional neural network of this embodiment. The method includes:
[0052] Step 101: Input the target sound source signal into the encoding model in the multi-task model to obtain the encoding result of the target sound source signal. Among them, the encoding model is used to encode the position information of the target sound source signal.
[0053] Among them, the sound source localization and recognition method based on the position-encoding convolutional neural network in this embodiment can be applied to fields such as underwater monitoring, security monitoring, medical monitoring, smart home, and urban intelligent management.
[0054] The sound source localization and recognition method based on the position-encoding convolutional neural network in this embodiment can be applied to different systems or devices, such as actuators. The actuator can be a smart terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, and a vehicle terminal, etc., or it can also be a server or the cloud, etc. This embodiment does not make specific limitations on this.
[0055] The target sound source signal is a sound event to be localized and recognized, which can be underwater sound, vehicle running sound, footsteps sound, or other sounds. It can be a single type of sound event or two overlapping sound events, etc. This embodiment does not make specific limitations on this.
[0056] It should be noted that the target sound source signal can be obtained by real-time acquisition based on a sound acquisition device or from the local storage of the sound acquisition device. This embodiment does not make specific limitations on the source of the target sound source signal.
[0057] Among them, the multi-task model is used to process the localization task and the recognition task of the target sound source signal; the encoding model is used to encode the position information of the target sound source signal.
[0058] Optionally, before localizing and recognizing the target sound source signal, the multi-task model needs to be trained. During the training stage of the multi-task model, a continuous audio of a certain duration is usually input as a sample sound source signal, and various sound events occur randomly within this duration. The multi-task model extracts the feature information of various sound events, and finally compares the output result with the label to complete the training and form an optimal multi-task model.
[0059] Optionally, after obtaining the target sound source signal, it can be directly input into the encoding model in the multi-task model; or the target sound source signal can be processed in one or more ways, such as performing preliminary feature extraction on the target sound source signal to obtain a logarithmic mel spectrogram feature vector and an intensity feature vector; then, input it into the encoding model in the multi-task model. This embodiment does not make specific limitations on this.
[0060] Optionally, the formula for inputting the target sound source signal into the encoding model in the multi-task model to obtain the encoding result of the target sound source signal is:
[0061]
[0062] where t ij is the encoding result of the j - th dimension feature of the i - th time series in the encoding result; a is the encoding coefficient; n is the total length of the time series of the target sound source signal. The total length of the time series and the feature dimension of the encoding result are consistent with the target sound source signal to form encoding features that are easy to input into the feature extraction model in the multi - task model.
[0063] Among them, the encoding coefficient can be set according to actual needs, such as a = 0.75; it can also be obtained by optimizing algorithms, and this embodiment does not make specific limitations on this.
[0064] By encoding the position information of the target sound source signal, the interference of the time position information in the encoding result can be eliminated.
[0065] Step 102: Input the target sound source signal and the encoding result into the feature extraction model in the multi - task model to obtain the feature vector of the target sound source signal;
[0066] Among them, the feature extraction model can be a convolutional neural network model, etc., and this embodiment does not make specific limitations on this.
[0067] The feature extraction model can be a multi - task feature extraction model, that is, a feature extraction model related to the localization task and a feature extraction model related to the recognition task; it can also be a single - task feature extraction model, that is, a model used to extract comprehensive features that can be used for both classification tasks and recognition tasks. This embodiment does not make specific limitations on the structure of the feature extraction model.
[0068] Optionally, after encoding the position information of the target sound source signal, the target sound source signal and the encoding result can be combined to form input information, and then the input information is input into the feature extraction model in the multi - task model to obtain the feature vector of the target sound source signal.
[0069] In this embodiment, by adding an encoding model to the multi - task model to encode the time position information of the occurrence of the sound event, and inputting the target sound source signal and the encoding result into the feature extraction model in the multi - task model, the position deviation can be eliminated when extracting features, so as to better extract the features of the sound events occurring at each time point within the same time period, thereby solving the problems such as poor robustness caused by the feature extraction deviation caused by the incomplete displacement invariance of the traditional convolutional neural network in the prior art; realizing more accurate extraction of the effective features of the sound event, being easy to promote in the actual application scenario, and further improving the accuracy of sound localization and recognition.
[0070] Step 103: Input the feature vector of the target sound source signal into the decoding model in the multi-task model to obtain the decoding result of the target sound source signal;
[0071] Optionally, after obtaining the feature vector of the target sound source signal, the feature vector can be input into the decoding model in the multi-task model to perform reverse decoding on the feature vector to obtain the decoding result.
[0072] Among them, the formula for inputting the feature vector of the target sound source signal into the decoding model in the multi-task model to obtain the decoding result of the target sound source signal is:
[0073]
[0074] Among them, F ij is the decoding result of the j-th dimension feature of the i-th time series in the decoding result; f ij is the j-th dimension feature of the i-th time series in the feature vector; a is the encoding coefficient; m is the total length of the time series of the feature vector.
[0075] Step 104: Input the decoding result of the target sound source signal into the localization and recognition model in the multi-task model to obtain the localization result and recognition result of the target sound source signal;
[0076] Among them, the multi-task model is trained based on the sample sound source signal and the corresponding reference localization result and reference recognition result.
[0077] Among them, the localization and recognition model includes a localization model and a recognition model;
[0078] The localization result includes the horizontal angle and pitch angle of the target sound source signal; the recognition result contains the category of the target sound source signal, such as the sound of vehicle running, the sound of dog barking, or the sound of cat meowing, etc.
[0079] Optionally, after obtaining the decoding result of the target sound source signal, the encoding result can be input into the localization model and the recognition model respectively for further feature extraction; then, the localization result and the recognition result are output through the fully connected layers of the localization model and the recognition model.
[0080] It should be noted that the output result of the localization model can include the sound source localization information of a single sound event that occurs without overlap or two sound events that occur with overlap of the target sound source signal; the output result of the recognition model also includes the sound source recognition information of a single sound event that occurs without overlap or two sound events that occur with overlap.
[0081] For example, if the target sound source signal only contains a single barking sound event, the recognition model outputs that the category of the target sound source signal is barking; the localization model outputs the horizontal angle and pitch angle at which the barking sound event occurs.
[0082] If the target sound source signal contains overlapping barking sound events and meowing sound events, the recognition model outputs that the category of the barking sound event in the target sound source signal is barking, and the category of the meowing sound event is meowing; the localization model outputs the horizontal angle and pitch angle at which the barking sound event occurs, and the horizontal angle and pitch angle at which the meowing sound event occurs.
[0083] Existing sound localization and recognition methods do not consider the poor ability of traditional convolutional neural networks to retain relative position information and the incomplete displacement invariance. Instead, they directly use them as feature extraction tools, resulting in position-biased information carried in the features extracted from sound events occurring at different time periods. As a result, traditional convolutional neural networks cannot effectively eliminate the interference of the time position information of sound events, and thus cannot deepen the convolutional neural network to achieve a better feature extraction effect. Eventually, it leads to a barrier in improving the accuracy of the training model and reduces the robustness of the convolutional neural network, facing problems such as lack of accuracy in the localization and recognition of short-term and similar events.
[0084] In this embodiment, feature extraction is jointly performed through the combination of an encoding model and a feature extraction model, effectively eliminating the interference of time position information in the feature vector and deeply mining the essential features that affect the localization task and recognition task from the target sound source signal. It can better complete the feature extraction of sound events, achieve stronger localization and recognition performance, and is more conducive to practical applications. Moreover, by using a multi-task model to learn the implicit relationship between the localization task and recognition task of the target sound source signal, the accuracy of the localization result and recognition result of the target sound source signal can be effectively improved.
[0085] On the one hand, in this embodiment, after encoding the position information of the target sound source signal, feature extraction is performed to eliminate the interference of time position information in the feature vector and deeply mine the essential features that affect the localization task and recognition task from the target sound source signal, thereby effectively improving the localization accuracy and recognition accuracy of the target sound source signal; on the other hand, a multi-task model is used to jointly learn the localization task and recognition task at the same time, fully considering the correlation and difference between the localization task and recognition task, and further improving the localization accuracy and recognition accuracy of the target sound source signal.
[0086] Based on the above embodiments, in this embodiment, the feature extraction model includes a first sub-feature extraction model and a second sub-feature extraction model, and the feature vector includes a first sub-feature vector and a second sub-feature vector; correspondingly, the step of inputting the target sound source signal and the encoding result into the feature extraction model in the multi-task model to obtain the feature vector of the target sound source signal includes: inputting the target sound source signal and the encoding result into the first sub-feature extraction model to obtain the first sub-feature vector of the target sound source signal, and inputting the target sound source signal and the encoding result into the second sub-feature extraction model to obtain the second sub-feature vector of the target sound source signal; wherein, the first sub-feature extraction model is used to extract features related to the localization result of the target sound source signal, and the second sub-feature extraction model is used to extract features related to the recognition result of the target sound source signal.
[0087] Among them, the feature extraction model is a two-way parallel convolutional neural network model; one way is the first sub-feature extraction model, which is used to extract features related to the localization result of the target sound source signal; the other way is the second sub-feature extraction model, which is used to extract features related to the recognition result of the target sound source signal. Each way is responsible for feature extraction of a single task of localization or recognition.
[0088] Among them, the structure of the first sub-feature extraction model is the same as that of the second sub-feature extraction model, but the internal parameters are different.
[0089] Since there is a certain correlation between the localization task and the recognition task to a certain extent, they are related tasks; a multi-task learning interaction module is added between the first sub-feature extraction model and the second sub-feature extraction model to facilitate the first sub-feature extraction model and the second sub-feature extraction model to mutually utilize the knowledge of the localization task and the recognition task to improve the performance of the feature extraction model.
[0090] Optionally, the input information formed by merging the target sound source signal and the encoding result is respectively input into the first sub-feature extraction model and the second sub-feature extraction model to respectively obtain the first sub-feature vector output by the first sub-feature extraction model and the second sub-feature vector output by the second sub-feature extraction model;
[0091] Then, the first sub-feature vector and the second sub-feature vector are respectively decoded, and the decoding result of the first sub-feature vector is input into the localization model to output the localization result of the target sound source signal; the decoding result of the second sub-feature vector is input into the recognition model to output the recognition result of the target sound source signal.
[0092] In this embodiment, the first sub-feature extraction model and the second sub-feature extraction model are used to extract features from the input information formed by the target sound source signal and the coding result in parallel, so as to extract the first sub-feature vector related to the positioning result of the target sound source signal and the second sub-feature vector related to the recognition result of the target sound source signal; furthermore, the first sub-feature vector contains more effective essential features for affecting the positioning task; the second sub-feature vector contains more effective essential features for affecting the recognition task, thereby making the positioning result and the recognition result more accurate.
[0093] Based on the above embodiment, in this embodiment, the feature extraction model includes at least one set of position information preservation module and pooling module; the position information preservation module includes a plurality of first convolution modules with different scales and a second convolution module; the plurality of first convolution modules with different scales are used to perform multi-scale feature extraction on the target sound source signal and the coding result to obtain a plurality of feature vectors of the target sound source signal with different scales; the second convolution module is used to fuse the feature vectors with different scales; the pooling module is used to perform a pooling operation on the fusion result.
[0094] Among them, the feature extraction model includes one or more sets of position information preservation module and pooling module, and the specific number of sets can be set according to actual needs.
[0095] As Figure 2 shown, the feature extraction model includes multiple sets of position information preservation module and pooling module; in each set, the output of the position information preservation module Blook is used as the input of the pooling module; in multiple sets, the output of the previous set of pooling module is used as the input of the next set of position information preservation module Blook.
[0096] Among them, the pooling module can be max pooling or average pooling Avgpool, and this embodiment does not make a specific limitation on this.
[0097] Each position information preservation module includes a plurality of first convolution modules with different scales and a second convolution module;
[0098] Among them, the number, layout position of the first convolution modules included in each position information preservation module, as well as the size and stride of the convolution kernel, etc. can all be set according to actual needs, and this embodiment does not make a specific limitation on this.
[0099] As Figure 3 shown is an exemplary structure of the information preservation module. The information preservation module includes a plurality of convolution modules, namely A1, A2, B1, B2, and B3 respectively; among them, the convolution kernel sizes of the A1, A2, B1, and B2 modules are all 3x3, and the strides are all 1; the convolution kernel size of the B3 module is 3x3, and the stride is 2;
[0100] Among them, by arranging A1, A2, B1, B2, and B3, multiple first convolution modules of different scales can be constructed.
[0101] Optionally, A1 and B1 are sequentially connected to form a first convolution module F1 of different scales; A1 and A2 are connected in parallel and then sequentially connected with B2 to form a first convolution module F2 of different scales; B3 is used as a first convolution module F3 of different scales;
[0102] C1 is used as the second convolution module, and the convolution kernel size of the C1 module is 1x1.
[0103] Optionally, the step of extracting features from input information formed by combining the target sound source signal and the encoding result includes:
[0104] First, the input information formed by combining the target sound source signal and the encoding result is input into the first convolution modules F1, F2 and F3 respectively;
[0105] Then, the feature vectors of different scales output by the first convolution modules F1, F2, and F3 and the input information are input into the second convolution module C1 for feature fusion;
[0106] Then, the fusion result is pooled through the pooling module;
[0107] Then, the above feature extraction process is repeated, and the pooling operation result is used as input information and input into the first convolution module of the next group to extract feature vectors of different scales until all groups have completed feature extraction;
[0108] Finally, the feature vector output by the last group of pooling modules is used as the feature vector of the target sound source signal.
[0109] In summary, the target sound source signal will output a high-level feature vector after passing through the encoding model and feature extraction model, which can effectively improve the positioning accuracy and recognition accuracy of the target sound source signal.
[0110] It should be noted that the first sub-feature extraction model and the second sub-feature extraction model have the same structure as the feature extraction model described above, and the feature extraction process is the same.
[0111] In this embodiment, a position information preservation module is used. On the one hand, it can mine feature vectors of multiple different scales to completely represent the feature vectors of each scale of the target sound source signal, and fuse the feature vectors of multiple different scales to represent the feature vectors related to the positioning task or recognition task of the target sound source signal, so as to improve the positioning accuracy and recognition accuracy of the target sound source signal; on the other hand, the entire position information preservation module realizes the fusion of diversified features without breaking the relative position relationship of the features of each layer of the convolution structure.
[0112] Based on the above embodiments, in this embodiment, the positioning and recognition model includes at least one set of parallel first Transformer models and second Transformer models; each set of the first Transformer models is used to locate each sound event of the target sound source signal; each set of the second Transformer models is used to recognize each sound event of the target sound source signal.
[0113] Among them, each set of the first Transformer models is used to locate a kind of sound event of the target sound source signal;
[0114] Each set of the second Transformer models is used to recognize a kind of sound event of the target sound source signal.
[0115] Optionally, the positioning and recognition model may include one or more sets of parallel first Transformer models and second Transformer models, and this embodiment does not make specific limitations on this.
[0116] For example, when locating and recognizing a single sound event, a positioning and recognition model including one set of the first Transformer models and the second Transformer models can be used to locate and recognize the single sound event in the target sound source signal.
[0117] When locating and recognizing two overlapping sound events, a positioning and recognition model including two sets of the first Transformer models and the second Transformer models can be used to locate and recognize the two sound events in the target sound source signal.
[0118] At this time, the positioning and recognition model altogether includes 4 parallel Transformer models, which are respectively used to further extract the positioning features and recognition features of sound event 1 in the target sound source signal, and the positioning features and recognition features of sound event 2; then, the positioning features and recognition features of sound event 1 and the positioning features and recognition features of sound event 2 are input into 4 fully connected neural networks for result output, and the output results are the positioning results and recognition results of sound event 1 and the positioning results and recognition results of sound event 2. Among them, the Transformer model adopts a structure of 8 heads and 2 layers, and the fully connected layer adopts a single-layer structure.
[0119] In this embodiment, the positioning and recognition model includes at least one set of parallel first Transformer models and second Transformer models, enabling the positioning and recognition model to not only locate and recognize a single sound event but also locate and recognize superimposed sound events, better coping with the positioning and recognition of overlapping sound events and having a wide range of applications.
[0120] Based on the above embodiments, before inputting the target sound source signal into the encoding model of the multi-task model to obtain the encoding result of the target sound source signal, this embodiment further includes: after performing preliminary data augmentation on the sample sound source signal, performing preliminary feature extraction to obtain the preliminary feature vector of the sample sound source signal; and / or, performing secondary data augmentation on some of the feature vectors in the preliminary feature vector of the sample sound source signal; training the multi-task model according to the preliminary feature vector of the sample sound source signal and / or the partially feature vectors after secondary data augmentation, as well as the reference positioning result and reference recognition result corresponding to the sample sound source signal.
[0121] Among them, the sample sound source signal can be an audio signal with a preset duration selected from the local storage of the sound acquisition device; the preset duration can be set according to the actual situation.
[0122] The sound acquisition device can be a FOA (First-Order Ambisonics) recording device, etc., and this embodiment does not make specific limitations in this regard; correspondingly, the format of the sample sound source signal can be an audio signal in FOA format;
[0123] Among them, the FOA recording device can collect audio signals from four channels: the x-direction, y-direction, z-direction, and omnidirectional audio channels.
[0124] Optionally, before inputting the target sound source signal into the multi-task model for positioning and recognition, the multi-task model needs to be trained;
[0125] In order to expand the training set, improve the robustness of the samples in different environments, and enable the trained multi-task model to adapt to background interference, with stronger robustness, generalization, and compatibility, when training the multi-task model in this embodiment, first, perform preliminary data augmentation on the sample sound source signal;
[0126] Among them, the preliminary data augmentation includes geometric transformation, noise augmentation, time shift augmentation, same-class augmentation, etc., and this embodiment does not make specific limitations in this regard.
[0127] Then, perform preliminary feature extraction on the sample sound source signal after preliminary data augmentation to obtain the preliminary feature vector of the sample sound source signal.
[0128] And / or, perform secondary data augmentation on some of the features in the preliminary feature vector;
[0129] Then, determine the sample input information based on the preliminary feature vector of the sample sound source signal and / or the partial feature vector after secondary data augmentation;
[0130] After sequentially inputting the sample input information into the encoding model, feature extraction model, decoding model, and localization and recognition model of the multi-task model, obtain the predicted localization result and predicted recognition result of the sample sound source signal;
[0131] Compare the predicted localization result of the sample sound source signal with the reference localization result to obtain the loss function corresponding to the localization model, and compare the predicted recognition result with the reference recognition result to obtain the loss function corresponding to the recognition model; jointly train the multi-task model with the loss functions corresponding to the recognition model and the localization model until the stop condition for training the multi-task model is met.
[0132] Among them, the way of jointly using the loss functions corresponding to the recognition model and the localization model can be to directly add the loss functions corresponding to the recognition model and the localization model as the overall loss function; it can also be to perform weighted addition of the loss functions corresponding to the recognition model and the localization model to balance the loss function of the entire multi-task model. This embodiment does not make specific limitations in this regard.
[0133] It should be noted that for a multi-task model that can be used to localize and recognize superimposed sound events, the superimposed training of the multi-task model can be completed in a permutation-invariant manner.
[0134] The trained multi-task model can be used for the localization and recognition of sound events. Since the sample sound source signal contains sample sound source signals in different directions and different environments after data augmentation, and effective features related to the recognition task and classification task of the sample sound source signal can be extracted by performing feature extraction on the sample sound source signal after data augmentation, the trained multi-task model can adapt to background interference, has stronger robustness, generalization ability, and compatibility, effectively improves the localization accuracy and recognition accuracy of the multi-task model, and can better handle the localization and recognition of overlapping sound events.
[0135] Based on the above embodiments, in this embodiment, the preliminary feature vector includes a logarithmic mel spectrogram feature vector and an intensity feature vector; correspondingly, performing secondary data augmentation on some of the feature vectors in the preliminary feature vector of the sample sound source signal includes: performing mel spectrogram data augmentation on the logarithmic mel spectrogram feature vector in the preliminary feature vector.
[0136] Among them, the preliminary feature vectors in this embodiment are not limited to log Mel spectrogram feature vectors and intensity feature vectors, and other feature vectors that can represent the target sound source signal can also be used.
[0137] Optionally, for the preliminary feature vectors of the sample sound source signal, obtain the preliminary feature vectors of the sample sound source signal;
[0138] Then, perform Mel spectrogram data augmentation SpecAugment on the log Mel spectrogram feature vectors in the preliminary feature vectors;
[0139] Among them, SpecAugment is a data augmentation method at the log Mel spectrogram level, which can convert the overfitting problem in model training into an underfitting problem, so as to alleviate the underfitting problem through large networks and long-term training strategies, and improve the recognition effect and localization effect of the model.
[0140] Finally, use the preliminary feature vectors of the sample sound source signal and / or the enhanced log Mel spectrogram feature vectors as sample input information to train the multi-task model.
[0141] Based on the above embodiments, the preliminary data augmentation in this embodiment includes rotating the sample sound source signal in one or more directions, and / or performing random superposition data augmentation on different categories of sample sound source signals.
[0142] Among them, the method of rotating the sample sound source signal in one or more directions is that the obtained sample sound source signal in FOA format contains audio signals in four channels: the x direction, the y direction, the z direction, and the omnidirectional audio channel W. The initial state of the audio signal in each channel is (α, β), where α is the azimuth angle and β is the elevation angle;
[0143] Rotate the audio signals in one or more channels in one or more directions to obtain an enhanced sample sound source signal;
[0144] Among them, the rotation azimuths include rotations in 15 directions such as (α - π / 2, β), (α - π / 2, -β), (α, -β), (α + π / 2, β), (α + π / 2, -β), (-α - π / 2, β), (-α - π / 2, -β), (-α, β), (-α, -β), (-α + π / 2, β), (-α + π / 2, -β), (-α + π, β), (-α + π, -β), (α + π, β), (α + π, -β);
[0145] Among them, the method of performing random superposition data augmentation on different categories of sample sound source signals is to perform random superposition data augmentation on the sample sound source signals in the training set with a certain probability. Among them, the probability can be set according to actual needs, such as the probability p = 0.6.
[0146] Optionally, in each batch, all sample sound source signals that have only a single event occurring throughout the duration are screened out, and the number is m; the number of remaining sample sound source signals X0 after screening is q. Let X1 be k samples among the m screened samples, X2 be k samples among the m screened samples excluding X1, X3 be v samples excluding the screened samples X1 and X2, and X be all samples in the entire batch after data augmentation, with the total number of samples being b. The relationship between data augmentation and the sample size is as follows:
[0147]
[0148] Among them, rand(α1,α2) is a floating-point number randomly generated between α1a1 and α2.
[0149] For example, if the total number of samples in each batch is 60 and the number of all sample sound source signals that have only a single event occurring in each batch is 32, then k = 15; X1 contains 15 samples, X2 contains 15 samples, and X3 contains 2 samples. Randomly superimpose the 15 samples in X1 and the 15 samples in X2 to obtain the superimposed samples and update X1; then, use the updated X1, as well as X0, X2, and X3 as the sample sound source signals after augmentation in this batch.
[0150] Different from existing sound event localization and detection methods, this embodiment uses azimuth augmentation and random combination superposition methods for the augmentation training of overlapping events, which can better localize and identify overlapping sound events.
[0151] As Figure 4 shown, it is a complete flow schematic diagram of the sound localization and recognition method based on the position-encoding convolutional neural network in this embodiment. The specific steps in the training stage include
[0152] Step 1, perform rotation data augmentation on the sample sound source signals in multiple azimuths;
[0153] Step 2, perform random superposition data augmentation on the sample sound source signals after rotation data augmentation;
[0154] Step 3, perform feature extraction on the sample sound source signals after random superposition data augmentation to obtain the number of Mel spectrogram feature vectors and intensity feature vectors, and perform logarithmic Mel spectrogram data augmentation on the Mel spectrogram feature vectors;
[0155] Step 4, combine the intensity feature vectors and the augmented logarithmic Mel spectrogram feature vectors to form input information, and after inputting the input information into the encoding model, input it into the feature extraction model to obtain the feature vectors of the sample sound source signals;
[0156] Step 5: Input the feature vector of the sample sound source signal into the decoding model for position information decoding to obtain the decoding result of the sample sound source signal; then, input the decoding result into the positioning model and the recognition model for feature extraction again, and output the prediction result through the fully connected layers of the positioning model and the recognition model; the prediction result includes the positioning result and the recognition result.
[0157] Step 6: Compare the prediction result with the label of the sample sound source signal and train the multi-task model.
[0158] In this embodiment, by randomly combining some non-overlapping events into overlapping events as the enhanced audio signal, performing primary feature extraction on the enhanced audio signal, and then using the encoding model and the feature extraction model to extract the feature vector that eliminates the position bias. And input the feature vector after eliminating the position bias into the positioning model and the recognition model, and compare the output result of the multi-task model with the label to complete the training of the multi-task model. The trained multi-task model can effectively complete the sound event localization and recognition tasks.
[0159] Next, the sound localization and recognition device based on the position-encoding convolutional neural network provided by the present invention will be described. The sound localization and recognition device based on the position-encoding convolutional neural network described below can be correspondingly referred to the sound localization and recognition method based on the position-encoding convolutional neural network described above.
[0160] As Figure 5 , this embodiment provides a sound localization and recognition device based on the position-encoding convolutional neural network. The device includes an encoding module 501, a feature extraction module 502, a decoding module 503, and a positioning and recognition module 504, where:
[0161] The encoding module 501 is configured to input the target sound source signal into the encoding model in the multi-task model to obtain the encoding result of the target sound source signal; wherein, the encoding model is used to perform position information encoding on the target sound source signal.
[0162] Optionally, before performing localization and recognition on the target sound source signal, it is necessary to train the multi-task model. In the multi-task model training stage, a continuous audio of a certain duration is usually input as the sample sound source signal, and various sound events occur randomly within this duration. The multi-task model is used to extract the feature information of various sound events, and finally the output result is compared with the label to complete the training and form the optimal multi-task model.
[0163] Optionally, after obtaining the target sound source signal, the target sound source signal can be directly input into the encoding model in the multi-task model; or one or more processes can be performed on the target sound source signal, such as performing preliminary feature extraction on the target sound source signal to obtain a logarithmic Mel spectrogram feature vector and an intensity feature vector; then, input it into the encoding model in the multi-task model. This embodiment does not make specific limitations on this.
[0164] Optionally, the formula for inputting the target sound source signal into the encoding model in the multi-task model to obtain the encoding result of the target sound source signal is:
[0165]
[0166] where t ij is the encoding result of the j-th dimension feature of the i-th time series in the encoding result; a is the encoding coefficient; n is the total length of the time series of the target sound source signal. The total length of the time series and the feature dimension of the encoding result are consistent with the target sound source signal to form an encoding feature that is easy to input into the feature extraction model in the multi-task model.
[0167] Among them, the encoding coefficient can be set according to actual needs, such as a = 0.75; it can also be obtained by optimizing the algorithm. This embodiment does not make specific limitations on this.
[0168] By performing position information encoding on the target sound source signal, the interference of time position information in the encoding result can be eliminated.
[0169] The feature extraction module 502 is configured to input the target sound source signal and the encoding result into the feature extraction model in the multi-task model to obtain the feature vector of the target sound source signal;
[0170] Among them, the feature extraction model can be a convolutional neural network model, etc. This embodiment does not make specific limitations on this.
[0171] The feature extraction model can be a multi-task feature extraction model, that is, a feature extraction model related to the classification task and a feature extraction model related to the recognition task are included; it can also be a single-task feature extraction model, that is, a comprehensive feature that can be used for both the classification task and the recognition task is extracted. This embodiment does not make specific limitations on the structure of the feature extraction model.
[0172] Optionally, after performing position information encoding on the target sound source signal, the target sound source signal and the encoding result can be combined to form input information, and then the input information is input into the feature extraction model in the multi-task model to obtain the feature vector of the target sound source signal.
[0173] The decoding module 503 is configured to input the feature vector of the target sound source signal into the decoding model in the multi-task model to obtain the decoding result of the target sound source signal;
[0174] Optionally, after obtaining the feature vector of the target sound source signal, the feature vector can be input into the decoding model in the multi-task model to perform reverse decoding on the feature vector to obtain the decoding result.
[0175] Among them, the formula for inputting the feature vector of the target sound source signal into the decoding model in the multi-task model to obtain the decoding result of the target sound source signal is:
[0176]
[0177] Among them, F ij is the decoding result of the j-th dimension feature of the i-th time series in the decoding result; f ij is the j-th dimension feature of the i-th time series in the feature vector; a is the encoding coefficient; m is the total length of the time series of the feature vector.
[0178] The positioning and recognition module 504 is configured to input the decoding result of the target sound source signal into the positioning and recognition model in the multi-task model to obtain the positioning result and recognition result of the target sound source signal;
[0179] Among them, the multi-task model is trained based on the sample sound source signal and the corresponding reference positioning result and reference recognition result.
[0180] The positioning result includes the horizontal angle and pitch angle of the target sound source signal; the recognition result includes the category of the target sound source signal, such as the sound of a vehicle running, the barking of a dog, or the meowing of a cat, etc.
[0181] Optionally, after obtaining the decoding result of the target sound source signal, the encoding result can be input into the positioning model and the recognition model respectively for re-feature extraction; then, the positioning result and the recognition result are output through the fully connected layers of the positioning model and the recognition model.
[0182] It should be noted that the output result of the positioning model can include the sound source positioning information of a single sound event that occurs without overlap or two sound events that occur with overlap of the target sound source signal; the output result of the recognition model also includes the sound source recognition information of a single sound event that occurs without overlap or two sound events that occur with overlap.
[0183] On the one hand, in this embodiment, after encoding the position information of the target sound source signal, feature extraction is performed to eliminate the interference of time position information in the feature vector, and the essential features affecting the localization task and the recognition task are deeply mined from the target sound source signal, thereby effectively improving the localization accuracy and recognition accuracy of the target sound source signal; on the other hand, a multi-task model is used to jointly learn the localization task and the recognition task at the same time, fully considering the correlation and difference between the localization task and the recognition task, and further improving the localization accuracy and recognition accuracy of the target sound source signal.
[0184] Based on the above embodiment, in this embodiment, the feature extraction model includes a first sub-feature extraction model and a second sub-feature extraction model, and the feature vector includes a first sub-feature vector and a second sub-feature vector; correspondingly, the feature extraction module is specifically configured to: input the target sound source signal and the encoding result into the first sub-feature extraction model to obtain the first sub-feature vector of the target sound source signal, and input the target sound source signal and the encoding result into the second sub-feature extraction model to obtain the second sub-feature vector of the target sound source signal; wherein, the first sub-feature extraction model is used to extract features related to the localization result of the target sound source signal, and the second sub-feature extraction model is used to extract features related to the recognition result of the target sound source signal.
[0185] Based on the above embodiment, in this embodiment, the feature extraction model includes at least one group of position information preservation module and pooling module; the position information preservation module includes a plurality of first convolution modules with different scales, and a second convolution module; the plurality of first convolution modules with different scales are used to perform multi-scale feature extraction on the target sound source signal and the encoding result to obtain a plurality of feature vectors of the target sound source signal with different scales; the second convolution module is used to fuse the plurality of feature vectors with different scales; the pooling module is used to perform a pooling operation on the fusion result.
[0186] Based on the above embodiments, in this embodiment, the localization and recognition model includes at least one group of parallel first Transformer models and second Transformer models; each group of the first Transformer models is used to localize each sound event of the target sound source signal; each group of the second Transformer models is used to recognize each sound event of the target sound source signal.
[0187] Based on the above embodiments, the present embodiment further includes a training module for: after performing preliminary data augmentation on the sample sound source signal, performing preliminary feature extraction to obtain a preliminary feature vector of the sample sound source signal; and / or, performing secondary data augmentation on some of the feature vectors in the preliminary feature vector of the sample sound source signal; training the multi-task model according to the preliminary feature vector of the sample sound source signal and / or the partially feature vectors after secondary data augmentation, as well as the reference localization result and reference recognition result corresponding to the sample sound source signal.
[0188] Based on the above embodiments, the preliminary feature vector in the present embodiment includes a logarithmic mel spectrogram feature vector and an intensity feature vector; correspondingly, the data augmentation module in the training module is used for: performing mel spectrogram data augmentation on the logarithmic mel spectrogram feature vector in the preliminary feature vector.
[0189] Based on the above embodiments, the preliminary data augmentation includes rotating the sample sound source signal in one or more directions, and / or randomly superimposing data augmentation on sample sound source signals of different categories.
[0190] Figure 6 An example of a schematic physical structure diagram of an electronic device is shown as Figure 6 As shown, the electronic device may include: a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 complete communication with each other through the communication bus 604. The processor 601 can call the logical instructions in the memory 603 to execute a sound localization and recognition method based on a position-encoding convolutional neural network. The method includes: inputting a target sound source signal into an encoding model in the multi-task model to obtain an encoding result of the target sound source signal; wherein, the encoding model is used for encoding the position information of the target sound source signal; inputting the target sound source signal and the encoding result into a feature extraction model in the multi-task model to obtain a feature vector of the target sound source signal; inputting the feature vector of the target sound source signal into a decoding model in the multi-task model to obtain a decoding result of the target sound source signal; inputting the decoding result of the target sound source signal into a localization and recognition model in the multi-task model to obtain a localization result and a recognition result of the target sound source signal; wherein, the multi-task model is trained based on a sample sound source signal and the reference localization result and reference recognition result corresponding to the sample sound source signal.
[0191] In addition, when the logical instructions in the above-mentioned memory 603 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0192] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the sound localization and recognition method based on a position-encoded convolutional neural network provided by the above-mentioned various methods. The method includes: inputting a target sound source signal into an encoding model in a multi-task model to obtain an encoding result of the target sound source signal; wherein, the encoding model is used to encode the position information of the target sound source signal; inputting the target sound source signal and the encoding result into a feature extraction model in the multi-task model to obtain a feature vector of the target sound source signal; inputting the feature vector of the target sound source signal into a decoding model in the multi-task model to obtain a decoding result of the target sound source signal; inputting the decoding result of the target sound source signal into a localization and recognition model in the multi-task model to obtain a localization result and an identification result of the target sound source signal; wherein, the multi-task model is trained based on sample sound source signals and the corresponding reference localization results and reference identification results.
[0193] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the sound localization and recognition method based on a position-encoded convolutional neural network provided by the above-mentioned various methods. The method includes: inputting a target sound source signal into an encoding model in a multi-task model to obtain an encoding result of the target sound source signal; wherein, the encoding model is used to encode the position information of the target sound source signal; inputting the target sound source signal and the encoding result into a feature extraction model in the multi-task model to obtain a feature vector of the target sound source signal; inputting the feature vector of the target sound source signal into a decoding model in the multi-task model to obtain a decoding result of the target sound source signal; inputting the decoding result of the target sound source signal into a localization and recognition model in the multi-task model to obtain a localization result and a recognition result of the target sound source signal; wherein, the multi-task model is trained based on a sample sound source signal and a reference localization result and a reference recognition result corresponding to the sample sound source signal.
[0194] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0196] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A sound localization and recognition method based on a position-encoded convolutional neural network, characterized in that Including: Input the target sound source signal into the encoding model in the multi-task model to obtain the encoding result of the target sound source signal; wherein, the encoding model is used to encode the position information of the target sound source signal; Input the target sound source signal and the encoding result into the feature extraction model in the multi-task model to obtain the feature vector of the target sound source signal; Input the feature vector of the target sound source signal into the decoding model in the multi-task model to obtain the decoding result of the target sound source signal; Input the decoding result of the target sound source signal into the positioning and recognition model in the multi-task model to obtain the positioning result and recognition result of the target sound source signal; Wherein, the multi-task model is trained based on the sample sound source signal and the corresponding reference positioning result and reference recognition result; The feature extraction model includes at least one set of position information preservation module and pooling module; The position information preservation module includes a plurality of first convolution modules with different scales and a second convolution module; The plurality of first convolution modules with different scales are used to perform multi-scale feature extraction on the target sound source signal and the encoding result to obtain a plurality of feature vectors of the target sound source signal with different scales; The second convolution module is used to fuse the feature vectors with different scales; The pooling module is used to perform a pooling operation on the fusion result; The position information preservation module includes a plurality of convolution modules, namely A1, A2, B1, B2, and B3; the kernel sizes of A1, A2, B1, and B2 are all 3x3, and the strides are all 1; the kernel size of B3 is 3x3, and the stride is 2; the plurality of first convolution modules with different scales include the first convolution module F1, the first convolution module F2, and the first convolution module F3. The first convolution module F1 is a first convolution module with different scales formed by sequentially connecting A1 and B1. The first convolution module F2 is a first convolution module with different scales formed by sequentially connecting A1 and A2 in parallel and then connecting B2. The first convolution module F3 is a first convolution module with different scales formed based on B3.
2. The method for sound localization and recognition based on a position-encoded convolutional neural network according to claim 1, wherein The feature extraction model includes a first sub-feature extraction model and a second sub-feature extraction model, and the feature vector includes a first sub-feature vector and a second sub-feature vector; Correspondingly, the step of inputting the target sound source signal and the encoding result into the feature extraction model in the multi-task model to obtain the feature vector of the target sound source signal includes: Input the target sound source signal and the encoding result into the first sub-feature extraction model to obtain the first sub-feature vector of the target sound source signal, and input the target sound source signal and the encoding result into the second sub-feature extraction model to obtain the second sub-feature vector of the target sound source signal; Wherein, the first sub-feature extraction model is used to extract features related to the positioning result of the target sound source signal, and the second sub-feature extraction model is used to extract features related to the recognition result of the target sound source signal.
3. The method for sound localization and recognition based on a position-encoded convolutional neural network according to any one of claims 1-2, characterized in that The positioning and recognition model includes at least one set of parallel first Transformer models and second Transformer models; Each set of the first Transformer models is used to position each sound event of the target sound source signal; Each set of the second Transformer models is used to recognize each sound event of the target sound source signal.
4. The sound localization and recognition method based on a position-encoded convolutional neural network according to any one of claims 1-2, characterized in that Before inputting the target sound source signal into the encoding model in the multi-task model to obtain the encoding result of the target sound source signal, it further includes: After performing preliminary data augmentation on the sample sound source signal, preliminary feature extraction is performed to obtain the preliminary feature vector of the sample sound source signal; And / or, perform secondary data augmentation on some of the feature vectors in the preliminary feature vector of the sample sound source signal; Train the multi-task model according to the preliminary feature vector of the sample sound source signal and / or the partial feature vectors after secondary data augmentation, as well as the reference positioning result and reference recognition result corresponding to the sample sound source signal.
5. The method for sound localization and recognition based on a position-encoding convolutional neural network according to claim 4, characterized in that, The preliminary feature vector includes a logarithmic Mel spectrogram feature vector and an intensity feature vector; Correspondingly, performing secondary data augmentation on some of the feature vectors in the preliminary feature vector of the sample sound source signal includes: Performing Mel spectrogram data augmentation on the logarithmic Mel spectrogram feature vector in the preliminary feature vector.
6. The method for sound localization and recognition based on a position-encoded convolutional neural network according to claim 4, characterized in that The preliminary data augmentation includes rotating the sample sound source signal in one or more directions, and / or randomly superimposing data augmentation on sample sound source signals of different categories.
7. An apparatus for sound localization and recognition based on a position-encoded convolutional neural network, characterized in that, It includes: An encoding module, configured to input the target sound source signal into the encoding model in the multi-task model to obtain the encoding result of the target sound source signal; wherein, the encoding model is used to perform position information encoding on the target sound source signal; A feature extraction module, configured to input the target sound source signal and the encoding result into the feature extraction model in the multi-task model to obtain the feature vector of the target sound source signal; A decoding module, configured to input the feature vector of the target sound source signal into the decoding model in the multi-task model to obtain the decoding result of the target sound source signal; A positioning and recognition module, configured to input the decoding result of the target sound source signal into the positioning and recognition model in the multi-task model to obtain the positioning result and recognition result of the target sound source signal; Among them, the multi-task model is trained based on the sample sound source signal and the reference positioning result and reference recognition result corresponding to the sample sound source signal; The feature extraction model includes at least one set of position information preservation modules and pooling modules; The position information preservation module includes a plurality of first convolution modules with different scales, and a second convolution module; The plurality of first convolution modules with different scales are used to perform multi-scale feature extraction on the target sound source signal and the encoding result to obtain a plurality of feature vectors of the target sound source signal with different scales; The second convolution module is used to fuse the feature vectors with different scales; The pooling module is used to perform a pooling operation on the fusion result; The position information retention module includes multiple convolutional modules, namely A1, A2, B1, B2, and B3 respectively; the convolutional kernels of A1, A2, B1, and B2 are all 3x3 in size, and the strides are all 1; the convolutional kernel of B3 is 3x3 in size, and the stride is 2; the multiple first convolutional modules with different scales include a first convolutional module F1, a first convolutional module F2, and a first convolutional module F3. The first convolutional module F1 is a first convolutional module with different scales formed by sequentially connecting A1 and B1. The first convolutional module F2 is a first convolutional module with different scales formed by sequentially connecting the parallel connection of A1 and A2 and B2. The first convolutional module F3 is a first convolutional module with different scales formed based on B3.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the sound localization and recognition method based on the position-encoded convolutional neural network according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the sound localization and recognition method based on the position-encoded convolutional neural network according to any one of claims 1 to 6.