Small sample sound event detection method based on frame level

By adopting a small sample sound event detection method based on the frame level in sound event detection, using PCEN voiceprint features and small sample sound event detection model, the problem of difficulty in processing small sample data is solved, and high-precision sound event detection is achieved.

CN119479653BActive Publication Date: 2025-05-13CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411527937.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-05-13
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process small sample data in sound event detection, especially in the case of difficulty in data collection and labeling, resulting in low detection accuracy.

Method used

Using a small sample sound event detection method based on the frame level, PCEN voiceprint features are extracted by preprocessing the target audio signal and loading it into the small sample sound event detection model for detection processing, and generating a predicted frame state to determine the sound event category and start and end time point.

Benefits of technology

High-precision detection of small sample sound events is achieved, and the accuracy and efficiency of sound event detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479653B_ABST
    Figure CN119479653B_ABST
Patent Text Reader

Abstract

The present invention relates to a small sample sound event detection method based on frame level. It includes: providing a target audio signal to be detected, and performing a first audio signal preprocessing on the target audio signal to extract a target signal feature set of the target audio signal; loading the extracted target signal feature set into a constructed small sample sound event detection model, so as to use the small sample sound event detection model to perform sound event detection on the target audio signal, so as to generate a predicted frame state corresponding to each audio frame in the current target signal PCEN voiceprint feature; based on the predicted frame state corresponding to all target signal PCEN voiceprint features, determine the sound event category contained in the target audio signal and the time start and end points corresponding to each sound event category. The present invention can effectively realize the detection of small sample sound events with high detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a sound event detection method, in particular to a small sample sound event detection method based on frame level. Background Art

[0002] The definition of sound event detection is: using smart devices to analyze audio signals and detect the types of sound events of interest in the audio signals and their corresponding start and end time points; specifically, sound event detection technology can be applied to scenarios related to monitoring and anomaly detection, such as: in healthcare supervision and daily monitoring, monitoring can be achieved by monitoring the crying of babies, the falling sounds of the elderly, etc.; in addition, in the field of industrial manufacturing, it can also be used to monitor abnormal working sounds of machinery, and can also monitor and feedback the sounds in smart home life; therefore, sound event detection technology is of great practical value.

[0003] In real life, sound events often occur simultaneously, and the duration of sound events is constantly changing, which makes data collection and annotation very difficult. In the early days of sound event detection, hidden Markov models, random forest regression, support vector machines and other methods were often used. With the continuous development of deep learning, more and more deep learning-based methods are indeed better than traditional methods in performance.

[0004] Deep learning methods require a large amount of data support, but the complexity of sound makes data collection difficult. Correctly labeling the sound event categories and the start and end times of their occurrence also requires enormous manpower and material resources; especially labeling the start and end times of the event, which requires judging and marking each moment in time in order to mark it more accurately, which is very labor-intensive.

[0005] In the sound event detection task, data that contains both event categories and the start and end times of the event is generally called strongly labeled data; data that only contains event categories but not the start and end times of the event is generally called weakly labeled data. Weakly labeled data is much easier to obtain than strongly labeled data; and in some special fields, the amount of data may be relatively small, making it difficult to reach the amount of data required for deep learning.

[0006] From the above description, it can be seen that in the process of sound event detection, especially for some special fields where data is difficult to obtain, strongly labeled data is difficult to obtain and the amount of data is small, which cannot meet the data volume required for deep learning. Therefore, how to realize the detection of small sample sound events is a difficult problem that needs to be solved urgently. Summary of the invention

[0007] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide a small sample sound event detection method based on frame level, which can effectively detect small sample sound events with high detection accuracy.

[0008] According to the technical solution provided by the present invention, a small sample sound event detection method based on frame level, the small sample sound event detection method comprises:

[0009] Providing a target audio signal to be detected, and performing a first audio signal preprocessing on the target audio signal to extract a target signal feature set of the target audio signal, wherein the target signal feature set includes a plurality of target signal PCEN voiceprint feature groups with time series characteristics, the target signal PCEN voiceprint feature groups include a plurality of target signal PCEN voiceprint features, and one target signal PCEN voiceprint feature corresponds to one audio frame of the target audio signal;

[0010] The extracted target signal feature set is loaded into the constructed small sample sound event detection model, so as to use the small sample sound event detection model to perform sound event detection on the target audio signal, wherein:

[0011] When performing sound event detection on the target audio signal, the target signal PCEN voiceprint feature groups are loaded into the small sample sound event detection model one by one according to the time series characteristics;

[0012] Performing detection processing on each target signal PCEN voiceprint feature group using a small sample sound event detection model, so as to generate a predicted frame state corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group after the detection processing;

[0013] Based on the predicted frame states corresponding to the PCEN voiceprint features of all target signals, the sound event categories contained in the target audio signal and the time start and end points corresponding to each sound event category are determined.

[0014] For each target signal PCEN voiceprint feature, we have:

[0015]

[0016] Among them, PCEN(t,f) is the voiceprint feature of the target signal PCEN, t is the time of the audio frame in the target audio signal corresponding to the voiceprint feature of the current target signal PCEN, f is the Mel spectrum response corresponding to the current audio frame in the target audio signal, E(t,f) is the input spectrum energy mean corresponding to the current audio frame in the target audio signal, ξ is a decimal constant factor, α is a gain factor, α∈(0,1), δ and γ are output range compression factors, and M(t,f) is a first-order IIR filter.

[0017] The small sample sound event detection model includes a feature extraction network, an information interaction network, a classifier network and a prediction frame splicing module connected in sequence, wherein:

[0018] When detecting and processing each target signal PCEN voiceprint feature group, a feature extraction network is used to extract features of the target signal PCEN voiceprint feature group to obtain target signal voiceprint extraction features of the target signal PCEN voiceprint feature group;

[0019] The information interaction network interacts with the target signal voiceprint extraction features and the multiple sound event reference features to generate information interaction first output information and information interaction second output information;

[0020] The classifier network fuses the first output information of the information interaction and the second output information of the information interaction, and determines the prediction frame state corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group through the prediction frame splicing module;

[0021] For all target signal PCEN voiceprint feature groups, based on the predicted frame state corresponding to each target signal PCEN voiceprint feature in each target signal PCEN voiceprint feature group, the predicted frame splicing module determines all sound event categories of the target audio signal and the time start and end points corresponding to each sound event category.

[0022] The information interaction network includes a frame-level input unit, a multi-head self-attention unit, and a secondary feature fusion unit, wherein:

[0023] The frame-level input unit includes a frame-level first input channel and a frame-level second input channel, wherein the target signal voiceprint extraction feature is loaded into the frame-level first input channel, and the multi-sound event reference feature is loaded into the frame-level second input channel;

[0024] Based on the target signal voiceprint extraction feature, the frame-level first input channel generates frame-level first channel output information, wherein the frame-level first channel output information directly jumps to output and forms the information interaction first output information of the information interaction network, and the frame-level first channel output information is also loaded to the Key end and the Value end of the multi-head self-attention unit;

[0025] Based on the multi-sound event reference features, the frame-level second input channel generates frame-level second channel output information, and the frame-level second channel output information is also loaded into the Query end of the multi-head self-attention unit;

[0026] The multi-head self-attention unit performs initial attention calculation on the frame-level first channel output information loaded via the Key end, the frame-level first channel output information loaded via the Value end, and the frame-level second channel output information loaded via the Query end, so as to generate attention feature information after the initial attention calculation;

[0027] Additively fusing the attention feature information with the frame-level first channel output information generated by the frame-level first input channel to form information interaction preliminary fusion information, and loading the information interaction preliminary fusion information into the secondary feature fusion unit;

[0028] The information interaction preliminary fusion information and the secondary feature fusion unit adopt residual connection to form information interaction secondary fusion information, and the information interaction secondary fusion information is used as the second output information of information interaction.

[0029] The classifier network includes a foreground and background classifier, an event classifier, and a classifier splicer, wherein:

[0030] The foreground and background classifier receives the first output information of the information interaction, and the event classifier receives the second output information of the information interaction;

[0031] The foreground and background classifiers, the event classifiers and the classification splicer are adapted and connected;

[0032] The first output information of the information interaction and the second output information of the information interaction are fused through the foreground and background classifier, the event classifier and the classification splicer to generate the prediction probability corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group;

[0033] For each prediction probability, the prediction frame splicing module generates a prediction frame state corresponding to each audio frame based on a preset prediction probability threshold.

[0034] When building a small sample sound event detection model, include:

[0035] Build a basic model for small sample sound event detection,

[0036] Constructing a pre-training data set, and using the pre-training data set to pre-train a basic model for small sample sound event detection, so as to generate a small sample sound event detection pre-trained model when the pre-training reaches a pre-training target state;

[0037] Based on the pre-trained model of small sample sound event detection, a first-stage fine-tuning model of small sample sound event detection and a second-stage fine-tuning model of small sample sound event detection are constructed;

[0038] A small sample fine-tuning stage training data set is constructed, and the small sample fine-tuning stage training data set is used to fine-tune the small sample sound event detection fine-tuning first stage model and the small sample sound event detection fine-tuning second stage model, wherein:

[0039] When performing fine-tuning training for each epoch, supervised training is first performed on the first stage model of small sample sound event detection fine-tuning, and then semi-supervised training is performed on the second stage model of small sample sound event detection fine-tuning;

[0040] When the fine-tuning training reaches the fine-tuning training target state, the small sample sound event detection fine-tuning second stage model configuration that reaches the fine-tuning training target state is used as the small sample sound event detection model.

[0041] The small sample sound event detection basic model includes a feature extraction basic network, an information interaction basic network and a classifier basic network, wherein:

[0042] The feature extraction basic network, the information interaction basic network and the classifier basic network are connected in sequence, and the classifier basic network includes a base class basic classifier and a foreground and background basic classifier;

[0043] When building a pre-training dataset, include:

[0044] Produce a large-scale audio data set, wherein the large-scale audio data set includes a plurality of basic audio signals, each of which is a strongly labeled audio;

[0045] Performing a second audio signal preprocessing on each basic audio signal to generate a basic audio feature set, wherein the basic audio feature set includes a plurality of basic signal PCEN voiceprint feature groups with time series characteristics and a basic audio category vector feature group corresponding to each basic signal PCEN voiceprint feature group;

[0046] A pre-training data set is constructed based on a basic audio feature set of a basic audio signal, wherein for any pre-training sample in the pre-training data set, it includes a basic signal PCEN voiceprint feature group and a basic audio category vector feature group corresponding to the basic signal PCEN voiceprint feature group;

[0047] Configure pre-training information and use the pre-training data set to train a basic model for small sample sound event detection, where:

[0048] The configured pre-training information includes the target number of pre-training rounds and the pre-training loss function;

[0049] When the number of training rounds of the small sample sound event detection basic model using the pre-training data set reaches the pre-training target number of rounds, the pre-training of the small sample sound event detection basic model reaches the pre-training target state, and a small sample sound event detection pre-trained model is generated based on the small sample sound event detection basic model.

[0050] The first-stage fine-tuning model for small sample sound event detection includes a first-stage feature extraction network, a first-stage information interaction network, and a first-stage classifier network, wherein:

[0051] The first-stage feature extraction network is generated by migrating the feature extraction basic network in the pre-trained model of small-sample sound event detection;

[0052] The first stage network of information interaction is generated by migrating the basic network of information interaction in the model pre-trained by small sample sound event detection;

[0053] The first-stage classifier network includes a first-stage first classifier, a first-stage second classifier, and a first-stage third classifier, wherein:

[0054] In the first stage, the first classifier is generated by transferring the base class basic classifier in the pre-trained model of small sample sound event detection;

[0055] The second classifier in the first stage is generated by transferring the basic foreground and background classifiers in the pre-trained model of small sample sound event detection;

[0056] The third classifier in the first stage is constructed based on the prototype feature extraction method;

[0057] The first classifier of the first stage is connected with the network of the first stage of feature extraction through adaptation;

[0058] The first-stage second classifier, the first-stage third classifier and the first-stage information interaction network adapter are connected.

[0059] The small sample fine-tuning stage training data set includes a fine-tuning first stage training data set and a fine-tuning second stage training data set, wherein:

[0060] The fine-tuning first stage model for small sample sound event detection is trained using the fine-tuning first stage training dataset;

[0061] The fine-tuning second stage model for small sample sound event detection is trained using the fine-tuning second stage training dataset;

[0062] The first stage of fine-tuning training training datasets include mixed domain datasets, single sound event support datasets, and multiple sound event support datasets;

[0063] When constructing the first stage training dataset for fine-tuning training, it includes:

[0064] Providing a registered audio set matching the target audio signal, wherein the registered audio set includes a plurality of registered audios, and each registered audio is a strongly marked audio;

[0065] Performing a third audio signal preprocessing on each registered audio to generate a registered audio feature set, wherein the registered audio feature set includes a plurality of registered audio PCEN voiceprint feature groups with time series features and a registered audio category vector feature group corresponding to each registered audio PCEN voiceprint feature,

[0066] Based on a registered audio PCEN voiceprint feature group and a corresponding registered audio category vector feature group, a registered audio training sample is formed;

[0067] Sampling a target number of pre-training samples in the pre-training dataset, and mixing the sampled pre-training samples with all registered audio training samples to form a mixed domain dataset;

[0068] Based on the registered audio feature set, a single sound event support dataset and multiple sound event support datasets are generated;

[0069] When performing model training on the small sample sound event detection fine-tuning first stage model, the corresponding first stage training samples in the mixed domain dataset, the single sound event support dataset, and the multiple sound event support datasets are respectively loaded into the small sample sound event detection fine-tuning first stage model to perform model training on the small sample sound event detection fine-tuning first stage model.

[0070] The small sample sound event detection fine-tuning second stage model includes a feature extraction second stage network, an information interaction second stage network and a classifier second stage network, wherein:

[0071] The second-stage network for feature extraction is generated by migrating the basic network for feature extraction in the pre-trained model for small-sample sound event detection.

[0072] The second-stage network of information interaction is generated by migrating the basic network of information interaction in the model pre-trained for small sample sound event detection;

[0073] The second stage network of the classifier includes a second stage first classifier and a second stage second classifier, wherein:

[0074] The first classifier in the second stage is generated by migrating the second classifier in the first stage;

[0075] The second classifier in the second stage is generated by migrating the third classifier in the first stage;

[0076] The first classifier of the second stage, the second classifier of the second stage and the second network adaptation connection of the information interaction stage;

[0077] When constructing the second stage training dataset for fine-tuning, it includes:

[0078] Acquire a reference audio signal matching the target audio signal, wherein the reference audio signal is a weakly labeled audio signal;

[0079] Performing a first audio signal preprocessing on the reference audio to generate a reference signal feature set of the reference audio signal, wherein the reference signal feature set includes a plurality of reference signal PCEN voiceprint feature groups with time series characteristics, and each reference signal PCEN voiceprint feature group is used as a fine-tuning second-stage training sample in a fine-tuning second-stage training data set;

[0080] When training the fine-tuning second-stage model for small-sample sound event detection, the time series characteristics are used to load the fine-tuning second-stage training samples one by one into the feature extraction second-stage network to extract reference signal voiceprint extraction features;

[0081] In the second stage of information interaction, the network interacts with the reference signal voiceprint extraction feature group and the multi-sound event training feature group, where:

[0082] The multi-sound event training feature group is generated by extracting and migrating the features of the first-stage training samples corresponding to the multiple sound event support data sets using the first-stage feature extraction network in the current epoch fine-tuning training;

[0083] The reference signal voiceprint extraction features are loaded into the frame-level first input channel of the second-stage network of information interaction;

[0084] The multi-sound event training features are loaded into the frame-level second input channel of the second-stage network of information interaction;

[0085] The information interaction first output information and the information interaction second output information of the information interaction second stage network are fused and output by the classifier second stage network to generate a predicted frame probability for each audio frame in the current fine-tuning second stage training sample;

[0086] Using the adaptive pseudo-label threshold thresh and the predicted frame probability of each audio frame, the fine-tuning second-stage training samples are configured as high-confidence pseudo-labels or low-confidence pseudo-labels, and the fine-tuning second-stage training samples corresponding to the high-confidence pseudo-labels are added to the fine-tuning training first-stage training data set to update the fine-tuning training first-stage training data set.

[0087] The advantages of the present invention are as follows: a first audio signal preprocessing is performed on the target audio signal to extract a target signal feature set of the target audio signal, and the extracted target signal feature set is loaded into a constructed small sample sound event detection model to use the small sample sound event detection model to perform sound event detection on the target audio signal, wherein the small sample sound event detection model is used to detect and process each target signal PCEN voiceprint feature to generate a predicted frame state corresponding to each audio frame in the current target signal PCEN voiceprint feature; based on the predicted frame states corresponding to all target signal PCEN voiceprint features, the sound event category contained in the target audio signal and the time start and end points corresponding to each sound event category are determined, thereby effectively realizing the detection of small sample sound events and improving the detection accuracy of sound events. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 The figure is a flow chart of an embodiment of a method for detecting sound events with a small sample size according to the present invention.

[0089] Figure 2 The present invention is a schematic diagram of an embodiment of performing a first audio signal preprocessing on a target audio signal.

[0090] Figure 3 The figure is a schematic diagram of an embodiment of the present invention when training a basic model for detecting sound events with small samples.

[0091] Figure 4 A schematic diagram of an embodiment of the information interaction network of the present invention.

[0092] Figure 5 A schematic diagram of an embodiment of a single sound event of the present invention.

[0093] Figure 6 A schematic diagram of an embodiment of multiple sound events of the present invention.

[0094] Figure 7 A schematic diagram of an embodiment of the temporal mask linear filter data enhancement of the present invention.

[0095] Figure 8 A schematic diagram of an embodiment of the fine-tuning training of the present invention. DETAILED DESCRIPTION

[0096] The present invention will be further described below in conjunction with specific drawings and embodiments.

[0097] In order to effectively detect small sample sound events and improve the detection accuracy of sound events, the present invention provides a small sample sound event detection method based on frame level. Specifically, the small sample sound event detection method includes:

[0098] Providing a target audio signal to be detected, and performing a first audio signal preprocessing on the target audio signal to extract a target signal feature set of the target audio signal, wherein the target signal feature set includes a plurality of target signal PCEN voiceprint feature groups with time series characteristics, the target signal PCEN voiceprint feature groups include a plurality of target signal PCEN voiceprint features, and one target signal PCEN voiceprint feature corresponds to one audio frame of the target audio signal;

[0099] The extracted target signal feature set is loaded into the constructed small sample sound event detection model, so as to use the small sample sound event detection model to perform sound event detection on the target audio signal, wherein:

[0100] When performing sound event detection on the target audio signal, the target signal PCEN voiceprint feature groups are loaded into the small sample sound event detection model one by one according to the time series characteristics;

[0101] Performing detection processing on each target signal PCEN voiceprint feature group using a small sample sound event detection model, so as to generate a predicted frame state corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group after the detection processing;

[0102] Based on the predicted frame states corresponding to the PCEN voiceprint features of all target signals, the sound event categories contained in the target audio signal and the time start and end points corresponding to each sound event category are determined.

[0103] Figure 1 A schematic diagram of an embodiment of small sample sound event detection is shown in . It can be seen from the figure that when performing sound event detection, it is necessary to provide a target audio signal to be detected and build a small sample sound event detection model. Thereafter, the small sample sound event detection model can be used to perform sound event detection on the target audio signal. It should be noted that the target audio signal is a signal that associates audio frequency with time. Generally, the horizontal axis of the target audio signal is time information, and the vertical axis of the target audio signal is audio frequency. Sound event detection on the target audio signal specifically refers to determining the sound event category contained in the target audio signal and the time start and end points corresponding to each sound event category. The sound event category generally refers to the category of the sound source or the category of the sound, such as the sound event category can be bird song, cat cry, dog bark, whistle and other categories; the time start and end points of the sound event category specifically refer to the start and end time and end time of the current sound event category in the target audio signal.

[0104] It is understandable that the target audio signal can be a sound signal obtained in a target scene, and the target scene for obtaining the target audio signal can be a scene for monitoring equipment in industrial production, or other application scenes that require detection of sound events. The method for obtaining the target audio signal can be consistent with the existing ones, such as being formed by picking up a microphone, etc. Due to the particularity of the target scene, there are fewer data samples available when constructing a small sample sound event detection model. Therefore, the small sample of the present invention specifically refers to fewer training data samples available when constructing a small sample sound event detection model. Specifically, the training data samples specifically refer to strongly labeled audio signals in the same target scene as the target audio signal.

[0105] In order to detect sound events on the target audio signal, it is necessary to perform a first audio signal preprocessing on the target audio signal. Specifically, the first audio signal preprocessing is performed on the target audio signal, that is, the target audio signal is preprocessed accordingly. By performing the first audio signal preprocessing on the target audio signal, a target signal feature set of the target audio signal can be obtained.

[0106] In one embodiment of the present invention, the target signal feature set includes several target signal PCEN (Per-channel energy normalization, PCEN) voiceprint feature groups, wherein the target signal PCEN voiceprint feature groups in the target signal feature set have a time series characteristic. The time series characteristic specifically refers to that all target signal PCEN voiceprint feature groups have a temporal correlation, and the timing of different target signal PCEN voiceprint feature groups is generally consistent with the sampling generation timing of the target audio signal. For example, based on the target audio signal in the starting time interval, the first target signal PCEN voiceprint feature group can be extracted; based on the target audio signal in the ending time interval, the last target signal PCEN voiceprint feature group can be extracted.

[0107] Specifically, each target signal PCEN voiceprint feature group includes several target signal PCEN voiceprint features, each target signal PCEN voiceprint feature corresponds to an audio frame of the target audio signal, and the correspondence between the target signal PCEN voiceprint feature and the audio frame can be referred to the following description of the method for generating the target signal PCEN voiceprint feature group. Generally, the number of target signal PCEN voiceprint features contained in each target signal PCEN voiceprint feature group is consistent, and each target signal PCEN voiceprint feature is k-dimensional data, and the situation of k-dimensional data will be described in detail below. The method of performing the first audio signal preprocessing on the target audio signal and obtaining the target signal feature set can refer to the following corresponding description.

[0108] When using a small sample sound event detection model to perform sound event detection on a target audio signal, the target signal PCEN voiceprint feature groups should be loaded into the small sample sound event detection model one by one according to the time series characteristics of different target signal PCEN voiceprint feature groups. As can be seen from the above description, the first time sequence target signal PCEN voiceprint feature group can be first loaded into the small sample sound event detection model, and thereafter, the corresponding target signal PCEN voiceprint feature groups are loaded into the small sample sound event detection model in turn until the last time sequence target signal PCEN voiceprint feature group is loaded into the small sample sound time detection model. At this point, all target signal PCEN voiceprint feature groups are loaded into the small sample sound event detection model one by one, and the small sample sound event detection model performs the same detection processing on each target signal PCEN voiceprint feature group. The detection processing process can refer to the corresponding description below.

[0109] From the above description, it can be seen that the target signal PCEN voiceprint feature group may include multiple target signal PCEN sonar features, and the multiple target signal PCEN voiceprint features correspond one-to-one to the multiple audio frames in the target audio signal. It can be seen that a corresponding target signal PCEN voiceprint feature group can be generated based on the corresponding number of audio frames in the target audio signal, and the number of target signal PCEN voiceprint features in the target signal PCEN voiceprint feature group is related to the first preprocessing method of the audio signal.

[0110] After a target signal PCEN voiceprint feature group is loaded into the small sample sound event detection model, the predicted frame state corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group can be determined after detection processing. It can be seen from the above description that one target signal PCEN voiceprint feature in the target signal PCEN voiceprint feature group corresponds to one audio frame. It can be understood that when multiple PCEN voiceprint features in a target signal PCEN voiceprint feature group correspond to multiple audio frames of the target audio signal, after detection processing, the predicted frame state of all corresponding audio frames of the current target signal PCEN voiceprint feature group can be determined. For example, when a target signal PCEN voiceprint feature group includes 431 target signal PCEN voiceprint features, at this time, corresponding to 431 audio frames in the target audio signal, after detection processing, the predicted frame state corresponding to the corresponding 431 audio frames can be determined.

[0111] It can be understood that when all the target signal PCEN voiceprint feature groups are loaded into the small sample sound event detection model, the frame prediction state of the audio frames corresponding to all the target signal PCEN voiceprint features can be obtained. At this time, the frame prediction state corresponding to all audio frames in the target audio signal can be obtained. Thereafter, based on the predicted frame state of all audio frames, the sound event category contained in the target audio signal and the time start and end points corresponding to each sound event category can be determined, thereby realizing the sound event detection of the target audio signal; among them, the small sample sound event detection model, frame prediction state, and the method and process of determining the category of the sound event contained in the target audio signal and the corresponding time start and end points will be described in detail below.

[0112] In one embodiment of the present invention, for each target signal PCEN voiceprint feature, there is:

[0113]

[0114] Among them, PCEN(t,f) is the voiceprint feature of the target signal PCEN, t is the time of the audio frame in the target audio signal corresponding to the voiceprint feature of the current target signal PCEN, f is the Mel spectrum response corresponding to the current audio frame in the target audio signal, E(t,f) is the input spectrum energy mean corresponding to the current audio frame in the target audio signal, ξ is a decimal constant factor, α is a gain factor, α∈(0,1), δ and γ are output range compression factors, and M(t,f) is a first-order IIR filter.

[0115] Specifically, since a target signal PCEN voiceprint feature can be generated based on an audio frame, the current audio frame in the above formula (1) is the audio frame in the target audio signal corresponding to the current target signal PCEN voiceprint feature.

[0116] It can be seen from the above description that in order to obtain the target signal PCEN voiceprint feature group, the first audio signal preprocessing should be performed on the target audio signal. Specifically, the first audio signal preprocessing performed on the target audio signal includes at least audio signal resampling processing, PCEN voiceprint feature conversion processing and PCEN voiceprint feature extraction processing, wherein the audio signal resampling processing specifically refers to sampling the target audio signal. When performing the audio signal resampling processing, the target audio signals of different sampling frequencies can be sampled into a uniform frequency to facilitate the subsequent PCEN voiceprint feature extraction processing; specifically, the sampling frequency of the target audio signal can be selected according to actual needs, such as a sampling frequency of 22050Hz. After the audio signal resampling processing is performed on the target audio signal, the target audio sampling signal can be generated.

[0117] After the target audio sampling signal is generated, a PCEN voiceprint feature conversion process is performed on the target audio sampling signal, so that a target signal spectrogram corresponding to the target audio signal can be obtained after the PCEN voiceprint feature conversion process is performed.

[0118] It should be noted that the method of executing the PCEN voiceprint feature conversion processing can be consistent with the existing PCEN voiceprint feature conversion method. A feasible PCEN voiceprint feature conversion processing method can be: the target audio sampling signal is sequentially subjected to frame processing, windowing processing, STFT (short-time Fourier transform) transformation processing, MEL filtering and PCEN voiceprint feature calculation. The above formula (1) is an embodiment of PCEN voiceprint feature calculation, wherein the corresponding processing processes of frame processing, windowing processing and STFT transformation processing can be consistent with the existing ones, such as the frame processing can adopt a 25ms frame method, the windowing processing can adopt a 0.5s Hanning window, the STFT transformation processing can adopt the existing commonly used short-time Fourier transform method, and the MEL filter can adopt the existing commonly used filtering method.

[0119] The above formula (1) shows a method for calculating the PCEN voiceprint feature, wherein the input spectrum energy mean E(t,f) corresponding to the current audio frame in the target audio signal can be obtained by calculating the mean of the audio spectrum of the target audio signal, and the method for calculating the mean of the audio spectrum of the target audio signal can be consistent with the existing method. The decimal constant factor ξ can generally be 1e-6, and the output range compression factor δ and the output range compression factor γ can have a corresponding value range of 0.3 to 0.7.

[0120] Furthermore, when calculating the PCEN voiceprint feature, a first-order IIR filter M(t,f) can be used to suppress irrelevant noise signals. The calculation method of the first-order IIR filter M(t,f) can be: M(t,f)=(E*φ T )(t,f)=sE(t,f)+(1-s)M(t-τ,f), where, φ T is a low-pass filter with a gain of 0 dB, s is a weight factor, ranging from 0 to 1, t is the current moment, and τ is the minimum time. Of course, other methods can also be used to implement the PCEN voiceprint feature conversion processing. The specific conversion processing method can be selected according to needs, and no examples will be given here.

[0121] From the above description, it can be known that after the PCEN voiceprint feature conversion processing, the target signal spectrogram corresponding to the target audio signal can be obtained. Therefore, for the target signal spectrogram corresponding to the target audio signal, the PCEN voiceprint feature corresponding to each audio frame can be obtained. Thereafter, the target signal spectrogram can be subjected to PCEN voiceprint feature extraction processing, so that after PCEN voiceprint feature extraction, the corresponding target signal PCEN voiceprint feature group can be obtained.

[0122] As can be seen from the above description, each target signal PCEN voiceprint feature is k-dimensional, where the k dimension here refers to the frequency dimension of the target signal PCEN voiceprint feature. As can be seen from the above description, the target signal PCEN voiceprint feature is calculated based on the Mel spectrum. After the target audio signal is framed and windowed in the time domain, each frame is subjected to a short-time Fourier transform (STFT) to convert it into a frequency domain representation. The result of the STFT transformation is a complex array, which represents the amplitude and phase of different frequency components. The frequency on the frequency axis is converted to the Mel scale to obtain the Mel spectrum, and then the number of Mel filters n_mels is specified as k. For example, k can be 128. At this time, the Mel spectrum of each frame will have 128 frequency dimensions. Since the PCEN voiceprint feature is a feature obtained by normalizing the Mel spectrum after filtering, that is, the frequency dimension of each target signal PCEN feature is the same as the Mel spectrum, which is 128 dimensions; when k is other values, please refer to the description here.

[0123] When performing PCEN voiceprint feature extraction processing, a feasible method may be: configuring a target signal sliding window, and using the target signal sliding window to slide the window on the target signal spectrogram. It can be seen that based on the PCEN voiceprint features of the audio frame corresponding to each target signal sliding window, a target signal PCEN voiceprint feature group can be formed; when sliding the window, there is an overlap between two adjacent target signal sliding windows. Generally, 86 frames of PCEN voiceprint features can be slid between two adjacent PCEN voiceprint extraction sample windows, which corresponds to a window shift of 1s. For example, for the first target signal sliding window and the second target signal sliding window, the starting point of the first target signal sliding window is the target signal PCEN voiceprint feature corresponding to the first audio frame, and the starting point of the second target signal sliding window should be the target signal PCEN voiceprint feature corresponding to the 86th audio frame. Other situations are similar and will not be explained one by one here.

[0124] Specifically, the audio frame corresponding to each target signal PCEN voiceprint feature group is related to the target signal sliding window configured above. From the above description, it can be seen that the number of target signal PCEN voiceprint features in all target signal PCEN voiceprint feature groups is consistent; in one embodiment of the present invention, each target signal PCEN voiceprint feature group may include 431 target signal PCEN voiceprint features, that is, each target signal PCEN voiceprint feature group corresponds to 431 audio frames in the target audio signal.

[0125] In one embodiment of the present invention, the small sample sound event detection model includes a feature extraction network, an information interaction network, a classifier network and a prediction frame splicing module connected in sequence, wherein:

[0126] When detecting and processing each target signal PCEN voiceprint feature group, a feature extraction network is used to extract features of the target signal PCEN voiceprint feature group to obtain target signal voiceprint extraction features of the target signal PCEN voiceprint feature group;

[0127] The information interaction network interacts with the target signal voiceprint extraction features and the multiple sound event reference features to generate information interaction first output information and information interaction second output information;

[0128] The classifier network fuses the first output information of the information interaction and the second output information of the information interaction, and determines the prediction frame state corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group through the prediction frame splicing module;

[0129] For all target signal PCEN voiceprint feature groups, based on the predicted frame state corresponding to each target signal PCEN voiceprint feature in each target signal PCEN voiceprint feature group, the predicted frame splicing module determines all sound event categories of the target audio signal and the time start and end points corresponding to each sound event category.

[0130] In order to detect sound events in the target audio signal, the small sample sound event detection model may include a feature extraction network, an information interaction network, a classifier network, and a prediction frame splicing module. The corresponding architecture of the small sample sound event detection model can be referenced. Figure 3 The basic model for small sample sound event detection shown, and Figure 8 The small sample sound event detection in the second stage fine-tunes the second stage model; it should be noted that the prediction frame splicing module is Figure 3 and Figure 8 The feature extraction network can be generated based on the CNN block. Figure 3 and Figure 8In the figure, an embodiment of a feature extraction network formed based on four CNN blocks connected in series is shown, that is, the 4×CNN in the figure is four CNN blocks connected in series, the CNN (Convolutional Neural Network) block can adopt the existing commonly used form, and the CNN block can adopt the default configuration of the existing commonly used CNN block. Of course, the feature extraction network can also adopt other forms, which shall be based on the satisfaction of feature extraction.

[0131] It can be seen from the above description that when performing sound event detection on the target audio signal, the target signal PCEN voiceprint feature group should be detected and processed. Specifically, for any target signal PCEN voiceprint feature group, the detection process includes sequentially performing corresponding processing through the feature extraction network, information interaction network, classifier network and prediction frame splicing module, wherein after the feature extraction network, the target signal voiceprint extraction feature can be extracted; the target signal voiceprint extraction feature and the multi-sound event reference feature are interacted through the information interaction network, and the information interaction first output information and the information interaction second output information can be obtained. Thereafter, the classifier network fuses the information interaction first output information and the information interaction second output information, and determines the prediction frame state corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group through the prediction frame splicing module, and all the sound event categories of the target audio signal and the time start and end points corresponding to each sound event category can be determined through the prediction frame splicing module.

[0132] In one embodiment of the present invention, the information interaction network includes a frame-level input unit, a multi-head self-attention unit, and a secondary feature fusion unit, wherein:

[0133] The frame-level input unit includes a frame-level first input channel and a frame-level second input channel, wherein the target signal voiceprint extraction feature is loaded into the frame-level first input channel, and the multi-sound event reference feature is loaded into the frame-level second input channel;

[0134] Based on the target signal voiceprint extraction feature, the frame-level first input channel generates frame-level first channel output information, wherein the frame-level first channel output information directly jumps to output and forms the information interaction first output information of the information interaction network, and the frame-level first channel output information is also loaded to the Key end and the Value end of the multi-head self-attention unit;

[0135] Based on the multi-sound event reference features, the frame-level second input channel generates frame-level second channel output information, and the frame-level second channel output information is also loaded into the Query end of the multi-head self-attention unit;

[0136] The multi-head self-attention unit performs initial attention calculation on the frame-level first channel output information loaded via the Key end, the frame-level first channel output information loaded via the Value end, and the frame-level second channel output information loaded via the Query end, so as to generate attention feature information after the initial attention calculation;

[0137] Additively fusing the attention feature information with the frame-level first channel output information generated by the frame-level first input channel to form information interaction preliminary fusion information, and loading the information interaction preliminary fusion information into the secondary feature fusion unit;

[0138] The information interaction preliminary fusion information and the secondary feature fusion unit adopt residual connection to form information interaction secondary fusion information, and the information interaction secondary fusion information is used as the second output information of information interaction.

[0139] Specifically, the information interaction network may include a frame-level input unit, a multi-head self-attention unit, and a secondary feature fusion unit, and the frame-level input unit, the multi-head self-attention unit, and the secondary feature fusion unit are adaptively connected. In a specific implementation, the frame-level input unit may include a frame-level first input channel and a frame-level second input channel. When performing sound event detection on a target audio signal, the target signal voiceprint feature extraction feature is loaded into the frame-level first input channel, and the multi-sound event reference feature is loaded into the frame-level second input channel; wherein, the frame-level first channel output information can be obtained through the frame-level first input channel, and the frame-level second channel output information can be obtained through the frame-level second input channel.

[0140] Figure 4 An embodiment of an information interaction network is shown in the figure, and an embodiment of a frame-level input unit is shown in the figure. In the figure, the first frame-level input channel may include a SED (Sound Event Detection) branch vector module, and the second frame-level input channel may include a FSBC (Foreground Sound Branch Classification) branch vector module and a POS (Postive) prototype module. The corresponding frame-level input processing performed by the first frame-level input channel and the second frame-level input channel can refer to the corresponding description below.

[0141] Figure 4 In the example, the multi-head self-attention unit has a Value end, a Key end, and a Query end. Specifically, the multi-head self-attention unit includes a first linear layer, a second linear layer, a third linear layer, a multi-head self-attention layer, and a first Dropout module. Figure 4, the linear layer corresponding to LN1 is the first linear layer, the linear layer corresponding to LN2 is the second linear layer, the linear layer corresponding to LN3 is the third linear layer, and the Dropout module corresponding to DP1 is the first Dropout module; wherein the Value end is formed by the input end of the first linear layer, the Key end is formed by the input end of the second linear layer, and the Query end can be formed by the input end of the third linear layer. The corresponding output ends of the first linear layer, the second linear layer, and the third linear layer are all connected to the input end of the multi-head self-attention layer, the output end of the multi-head self-attention layer is connected to the first Dropout module, and the output end of the first Dropout module is connected to the first splicer of the interactive network. Figure 4 In the figure, Cat1 is the first interactive network stitcher. The first interactive network stitcher can be used to realize additive fusion in the time dimension, that is, the preliminary fusion information of information interaction can be obtained through the first interactive network stitcher.

[0142] Figure 4 In the figure, the secondary feature fusion unit includes a layer normalization module, a first fully connected layer, a GELU function, a second Dropout module, a second fully connected layer and a second Dropout module connected in sequence. In the figure, the fully connected layer corresponding to QN1 is the first fully connected layer, the fully connected layer corresponding to QN2 is the second fully connected layer, the Dropout module corresponding to DP2 is the second Dropout module, and the Dropout module corresponding to DP3 is the third Dropout module. The layer normalization module is connected to the output end of the first splicer of the interactive network, the output end of the third Dropout module is connected to an input end of the second splicer of the interactive network, and the input of the layer normalization module is also connected to the other input end of the second splicer of the interactive network, so that the secondary feature fusion unit can form a residual connection form.

[0143] Figure 4 In the figure, Jout1 is the first output end of the information interaction network, Jout2 is the second output end of the information interaction network, the first output information of the information interaction can be obtained through the first output end of the information interaction network, and the second output information of the information interaction can be obtained through the second output end of the information interaction network.

[0144] Specifically, for the frame-level input unit, the target signal voiceprint extraction features extracted by the feature extraction network and the multi-sound event reference features are used as inputs respectively, and the two outputs of the frame-level input unit are: s ,z p ∈R h×w×c , where w and h are the width and height of the corresponding input feature map, corresponding to the time axis and frequency axis of the target audio signal, respectively, c is the number of channels of the feature map, and z sis the feature map output after the multiple sound event reference features are processed by the SED branch vector module, z p It is the category prototype window output after the target signal voiceprint extraction features are processed by the FBSC branch vector module and the POS prototype module.

[0145] It should be noted that the SED branch vector module copies the reference features of multiple sound events so that the reference features of multiple sound events can be loaded into the Value end and the Key end of the multi-head self-attention unit at the same time, that is, the feature map z s It should be consistent with the reference features of multiple sound events. The FBSC branch vector module is used to perform frame-level label prediction on the target signal voiceprint extraction features and obtain the feature map z t , that is, the feature map z t The label category that represents the foreground and background corresponding to the target signal voiceprint extraction feature. In the label category, when the label value is 0, it represents that the current label is the background class, and when the label value is 1, it represents that the current label is the foreground class.

[0146] In the specific implementation, the prototype extraction method is used to extract the feature map z t Extract the category prototype, and then restore it in the time dimension through the repeat operation, such as Figure 4 As shown, the feature map z is realized through the POS prototype module t Prototype feature extraction is performed, and then a repeat operation is performed to restore the time dimension. After that, the restored category prototype window z p Passed to the Query end, specifically, when the target signal voiceprint extraction features are restored in the time dimension, there are:

[0147]

[0148] in, is the i-th label value of the audio frame belonging to the foreground category in the target signal voiceprint extraction feature, and η is the feature map z t The total number of audio frames belonging to the foreground category in , is the single-frame foreground category prototype, z p After the repeat operation, it is combined with the feature map z t Category prototype windows aligned in the temporal dimension.

[0149] Specifically, the feature map z t The value of each element in is 0 or 1. The target signal voiceprint extraction feature is obtained by extracting the target signal PCEN voiceprint feature group through the feature extraction network. Since downsampling is required, the target signal voiceprint extraction feature is generated by the FBSC branch vector module to generate a feature map z t , the feature map z tThe dimension of the target signal PCEN voiceprint feature group is generally lower than that of the feature map z t The number of elements in the target signal PCEN voiceprint feature group is lower than the number of target signal PCEN voiceprint features in the target signal PCEN voiceprint feature group), and after the above repeat operation, the category prototype window z p The dimension of is consistent with the dimension of the target signal PCEN voiceprint feature group; in addition, the feature map z t , category prototype window z p The corresponding dimension of each element in can be 128 dimensions, such as the feature map z t An element within Indicates that the 128-dimensional data of the current element are

[0150] It is understandable that the feature map z t The corresponding elements The situation is specifically related to the feature extraction network for the processed target signal PCEN voiceprint feature group. In the above description, the prototype refers to the extracted event feature that can represent the members of the event category.

[0151] Specifically, the total number of frames n belonging to the foreground category audio frame is the number of frames of the foreground category audio frame corresponding to a target signal sliding window. Therefore, the total number of frames n is related to the target audio signal and the corresponding position of the target signal sliding window. As can be seen from the above description, the category prototype window z p The dimension should be n, n can be 431; at this time, the category prototype window z p The number of elements in is 431.

[0152] When performing self-attention, the feature map z s As Vlue and Key respectively, the category prototype window z p As Query, according to the category prototype window z p The category prototype pair feature map z in s The attention calculation of the corresponding category sound events in the feature map z is increased s The attention weights for events in the foreground category are calculated as follows:

[0153] First, the first linear layer and the second linear layer are used to transform the feature map z s Perform high-dimensional linear mapping and use the third linear layer to map the category prototype window z p Perform high-dimensional feature mapping; then, use the multi-head self-attention unit to perform preliminary attention calculation, and add the first Dropout module to prevent the model from overfitting, then we have:

[0154]

[0155] Among them, Wi Q , W i K , W i V are the input mapping matrices of Q(Query), K(Key), and V(Value) in each head in the multi-head self-attention layer, l is the total number of heads, Attention is the attention calculation of the multi-head self-attention layer, and W O is the output matrix of the multi-head attention layer, and its output dimension is consistent with the input data channel dimension.

[0156] Specifically, the input mapping matrix and the output matrix are consistent with the existing ones. The total number of head heads l can generally be 8. Of course, other numbers can also be selected according to actual needs, based on the ability to meet actual small sample sound time detection.

[0157] After the multi-head self-attention layer completes the attention calculation, the feature map z is connected using the residual connection method. s Additive fusion is performed with the output attention feature to supplement the attention content; then, the output attention is sent to the secondary feature fusion unit in turn, and the GELU activation function is added to the secondary feature fusion unit for activation. At the same time, the second Dropout module is added after the first fully connected layer, and the third Dropout module is added after the second fully connected layer to prevent the model from overfitting during training; finally, the residual connection method is used to fuse the input content with the output content. The mathematical expression of the GELU activation function can be:

[0158]

[0159] Wherein, x is the feature map input to the GELU activation function.

[0160] At the output of the information interaction network, in the SED branch with the SED branch vector module, the feature map z is connected by skip connection. s Output directly to retain the original frame-level event detection task. In the foreground and background classification branch containing the FBSC branch vector module, the branch interaction information after the self-attention mechanism is output, and the foreground and background classifier in the classifier network is used for frame-by-frame classification. The situation of the foreground and background classifier can refer to the corresponding description below.

[0161] In one embodiment of the present invention, the classifier network includes a foreground and background classifier, an event classifier, and a classifier splicer, wherein:

[0162] The foreground and background classifier receives the first output information of the information interaction, and the event classifier receives the second output information of the information interaction;

[0163] The foreground and background classifiers, the event classifiers and the classification splicer are adapted and connected;

[0164] The first output information of the information interaction and the second output information of the information interaction are fused through the foreground and background classifier, the event classifier and the classification splicer to generate the prediction probability corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group;

[0165] For each prediction probability, the prediction frame splicing module generates a prediction frame state corresponding to each audio frame based on a preset prediction probability threshold.

[0166] Specifically, the foreground and background classifiers and the event classifiers can both adopt existing commonly used classifier forms, wherein the foreground and background classifiers can classify and process the first output information of the information interaction, and the event classifier can classify and process the second output of the information interaction. Thereafter, they can be fused through the classifier splicer, that is, the classifier network can be used to realize the fusion of the first output information of the information interaction and the second output information of the information interaction, and after the fusion, the prediction probability corresponding to each target signal PCEN voiceprint feature in the target signal PCEN voiceprint feature group is generated.

[0167] It should be noted that the prediction probability can be a prediction value of 0 to 1. After obtaining the prediction probability corresponding to each target signal PCEN voiceprint feature, for each prediction probability, the prediction frame state corresponding to each target signal PCEN voiceprint feature can be generated based on the preset prediction probability threshold. For example, the prediction probability threshold can be set to 0.5. When the prediction probability is not less than 0.5, the prediction frame state corresponding to each audio frame can be configured as "1", otherwise, the corresponding prediction frame state is configured as "0"; of course, the prediction probability threshold can also be set to other values, which can be selected according to needs; based on the prediction probability and the prediction probability threshold, the method and process of generating the corresponding prediction frame state can refer to the description here, and no examples are given here one by one.

[0168] It should be noted that after all the target signal PCEN voiceprint feature groups are loaded into the small sample sound event detection model, the prediction frame states corresponding to all the target signal PCEN voiceprint features can be obtained. After that, the prediction frame splicing module can use the median filtering method to splice the output prediction frames and then perform median filtering. Specifically, the prediction frame splicing module can restore the output results according to the sliding window segmentation order of the target signal spectrogram by the target signal sliding window. The output result here is the prediction frame state corresponding to all the target signal PCEN voiceprint feature groups. It can be seen from the above description that the target audio signal is divided into multiple small segments after the first preprocessing. Since the target audio signal is continuous, the prediction results are re-spliced ​​to restore continuity after segmentation prediction, that is, through sliding restoration and median filtering, it can be ensured that the prediction results are smoother and more continuous.

[0169] Specifically, the middle position of each window is taken according to the middle area alignment method to output the result. The restoration is to splice the prediction results in the target signal sliding window according to the original sliding order. For example, for each target signal sliding window, the middle part of the target signal sliding window is taken as the prediction frame state corresponding to the current target signal sliding window. For example, a target signal sliding window corresponds to an audio frame in a 5s time range (that is, 431 audio frames). At this time, the prediction frame state corresponding to the middle part of 1s (corresponding to 86 frames) in the 5s time range can be taken as the prediction frame state corresponding to the current target signal sliding window. In specific implementation, when median filtering is used, the size of the filter kernel can be 5.

[0170] From the above description, it can be seen that when performing sound event detection, it is based on the target signal PCEN voiceprint feature group, and each target signal PCEN voiceprint feature group is generated based on the target signal sliding window, and for two adjacent target signal PCEN voiceprint feature groups, the corresponding target signal sliding windows overlap. It can be seen that the prediction information cannot be directly located to the corresponding time point of the sound event in the target audio signal based on the information of the target signal sliding window. Therefore, the above-mentioned restoration operation is required, that is, the start and end time of the sound event in the target audio signal can be accurately located through the restoration operation.

[0171] In one embodiment of the present invention, when constructing a small sample sound event detection model, it includes:

[0172] Build a basic model for small sample sound event detection,

[0173] Constructing a pre-training data set, and using the pre-training data set to pre-train a basic model for small sample sound event detection, so as to generate a small sample sound event detection pre-trained model when the pre-training reaches a pre-training target state;

[0174] Based on the pre-trained model of small sample sound event detection, a first-stage fine-tuning model of small sample sound event detection and a second-stage fine-tuning model of small sample sound event detection are constructed;

[0175] A small sample fine-tuning stage training data set is constructed, and the small sample fine-tuning stage training data set is used to fine-tune the small sample sound event detection fine-tuning first stage model and the small sample sound event detection fine-tuning second stage model, wherein:

[0176] When performing fine-tuning training for each epoch, supervised training is first performed on the first stage model of small sample sound event detection fine-tuning, and then semi-supervised training is performed on the second stage model of small sample sound event detection fine-tuning;

[0177] When the fine-tuning training reaches the fine-tuning training target state, the small sample sound event detection fine-tuning second stage model configuration that reaches the fine-tuning training target state is used as the small sample sound event detection model.

[0178] In order to construct the above-mentioned small sample sound event detection model and perform effective sound event detection on the target audio signal, in the specific implementation, it is necessary to first construct a small sample sound event detection basic model, pre-train the small sample sound event detection basic model through a pre-training data set, and generate a small sample sound event detection pre-trained model when the pre-training reaches the pre-training target state. In addition, based on the small sample sound event detection pre-trained model, a small sample sound event detection fine-tuning first stage model and a small sample sound event detection fine-tuning second stage model can be constructed, and the small sample sound event detection fine-tuning first stage model and the small sample sound event detection fine-tuning second stage model are fine-tuned and trained through a small sample fine-tuning stage training data set, and the above-mentioned small sample sound event detection model can be finally constructed. The following is an example of the method and process of specifically constructing a small sample sound event detection model.

[0179] In one embodiment of the present invention, the small sample sound event detection basic model includes a feature extraction basic network, an information interaction basic network and a classifier basic network, wherein:

[0180] The feature extraction basic network, the information interaction basic network and the classifier basic network are connected in sequence, and the classifier basic network includes a base class basic classifier and a foreground and background basic classifier;

[0181] When building a pre-training dataset, include:

[0182] Creating a large-scale audio data set, wherein the large-scale audio data set includes a number of basic audio signals, each of which is a strongly labeled audio signal;

[0183] Performing a second audio signal preprocessing on each basic audio signal to generate a basic audio feature set, wherein the basic audio feature set includes a plurality of basic signal PCEN voiceprint feature groups with time series characteristics and a basic audio category vector feature group corresponding to each basic signal PCEN voiceprint feature group;

[0184] A pre-training data set is constructed based on a basic audio feature set of a basic audio signal, wherein for any pre-training sample in the pre-training data set, it includes a basic signal PCEN voiceprint feature group and a basic audio category vector feature group corresponding to the basic signal PCEN voiceprint feature group;

[0185] Configure pre-training information and use the pre-training data set to train a basic model for small sample sound event detection, where:

[0186] The configured pre-training information includes the target number of pre-training rounds and the pre-training loss function;

[0187] When the number of training rounds of the small sample sound event detection basic model using the pre-training data set reaches the pre-training target number of rounds, the pre-training of the small sample sound event detection basic model reaches the pre-training target state, and a small sample sound event detection pre-trained model is generated based on the small sample sound event detection basic model.

[0188] Figure 3 An embodiment of a basic model for detecting small sample sound events is shown in FIG. Figure 3 It can be seen that the feature extraction basic network should adopt the same form as the feature extraction network in the above small sample sound event detection model, that is, the form of 4 CNN modules connected in series. Similarly, the information interaction basic network should adopt the same form as the information interaction network in the above small sample sound event detection model. The situation of the information interaction basic network can be referred to Figure 4 And the description of the above information interaction network. The classifier basic network is different from the above classifier network. Figure 3 In the figure, the classifier corresponding to Decoder1 is the base class basic classifier, and the classifier corresponding to Decoder2 is the foreground and background basic classifier.

[0189] After building a basic model for small sample sound event detection, you need to build a corresponding pre-training dataset, which includes several pre-training samples. When building a pre-training dataset, you need to first build a large-scale audio dataset, which includes several basic audio signals. Each basic audio signal is a strongly labeled audio. According to the description of the strongly labeled audio, each basic audio signal has a label for the sound event and the start and end time points of each sound event.

[0190] In specific implementation, a large-scale audio dataset can be constructed based on common speech datasets such as the AudioSet dataset and the Dcase 2023Task5 dataset. For example, a large-scale audio dataset can be constructed based on the AudioSet dataset. At this time, the large-scale audio dataset may include all or part of the audio data in the AudioSet dataset. The number of basic audio signals in the large-scale audio dataset can be selected as needed to meet the pre-training requirements of the basic model for small sample sound event detection.

[0191] It should be noted that the pre-training of the small sample sound event detection basic model is mainly to enable the pre-trained small sample sound event detection model to have the ability to extract features and interact with audio signals; from this, it can be seen that when a small sample sound event detection model is constructed based on the small sample sound event detection basic model, the small sample sound event detection model can also have the corresponding feature extraction and information interaction capabilities.

[0192] In order to form a pre-training sample, it is necessary to perform a second audio signal preprocessing on each basic audio signal, so that after performing the second audio signal preprocessing, the basic signal PCEN voiceprint feature group of the basic audio signal can be obtained, wherein the method of extracting the basic signal PCEN voiceprint feature group of the basic audio signal can refer to the description of performing the first audio signal preprocessing to obtain the target signal PCEN voiceprint feature group, which will not be repeated here. Different from the above-mentioned first audio signal preprocessing on the target signal, when performing the second audio signal preprocessing on the basic audio signal, the basic audio category vector feature group of the basic audio signal can also be obtained.

[0193] During specific implementation, when the second audio signal preprocessing is performed on the basic audio signal and a basic audio category feature vector group is obtained, a feasible solution is: the second audio signal preprocessing may include audio signal resampling processing, PCEN voiceprint feature conversion processing and PCEN voiceprint feature extraction processing. Specifically, the corresponding processes of audio signal resampling processing and PCEN voiceprint feature conversion processing can refer to the above-mentioned corresponding processing instructions for the target audio signal, which will not be repeated here. The PCEN voiceprint feature extraction processing of the basic audio signal can be consistent with the PCEN voiceprint feature extraction processing of the target audio signal. It can be seen from the above description that when the PCEN voiceprint feature extraction processing is performed on the basic audio signal, a basic audio sliding window needs to be constructed to use the basic audio sliding window to segment the basic signal spectrogram. The basic audio sliding window can refer to the description of the above-mentioned target signal sliding window. Generally, the basic audio sliding window and the target signal sliding window should use the same window setting, such as the same window length.

[0194] It can be seen from the above description that the basic audio signal is a strongly labeled audio. When the sound event is labeled, the labeling method is the start and end time of the sound event; and after the above-mentioned second processing of the audio signal is performed, the basic signal PCEN voiceprint feature corresponding to each audio frame can be obtained. Therefore, the labeling time corresponding to the basic signal PCEN voiceprint feature in the basic audio sliding window is converted according to the frame-level labeling, and the audio frame position of the sound event can be obtained, wherein the frame-level labeling conversion is to convert the labeled time into the corresponding audio frame position. Thereafter, according to the labeled labels, the labeled category vectors in each basic audio sliding window are extracted. Specifically, the labeled category vectors are extracted, that is, the sound events in each basic audio sliding window are extracted, and the start and end audio frames of each sound event are determined. At this time, a basic audio category vector feature group can be obtained. The method of performing the second audio signal preprocessing on the basic audio signal can refer to. Figure 2 illustrate; Figure 2 In the method, the basic signal spectrogram is preprocessed by framing operation, and the frame-level data of events and non-events are annotated, where non-events are annotated as "class 0" (background class) to improve the generalization ability of the model. Subsequently, the annotated frame-level data is segmented using the basic audio sliding window. The second preprocessing method of the audio signal can also adopt other commonly used forms, which can be selected according to needs, so as to obtain the basic signal PCEN voiceprint feature group and the corresponding basic audio category vector feature group.

[0195] In a specific implementation, the basic signal PCEN voiceprint feature group includes a number of basic signal PCEN voiceprint features, wherein the number of basic signal PCEN voiceprint features in a basic signal PCEN voiceprint feature group is the same as the number of target signal PCEN voiceprint features in a target signal PCEN voiceprint feature group. In addition, the number of basic audio category vector features in the basic audio category vector feature group is consistent with the number of basic signal PCEN voiceprint features in the basic signal PCEN voiceprint feature group, and the dimension of the basic signal PCEN voiceprint feature is also consistent with the dimension of the basic audio category vector feature.

[0196] When a basic audio signal includes multiple sound event categories, foreground and background classification is required when extracting the basic audio category vector features. Specifically, when performing foreground and background classification, in a basic audio sliding window, the target category selection method can be used to perform adaptive foreground / background category division of the labeled data. One feasible method is: first, according to the sound event categories corresponding to different processing times in the current basic audio sliding window, each sound event is selected as the foreground category in turn. t ; Then, according to the foreground class c tThe basic audio sliding window is annotated with foreground / background categories frame by frame; finally, the PCEN frame-level units belonging to the background class are erased according to the annotated content. The basic audio category vector feature expression is as follows:

[0197]

[0198] in, is the k-dimensional basic audio category vector feature, y i is the event category label corresponding to each basic signal PCEN voiceprint feature, c t is the selected foreground class; in specific implementation, k can be 128. Therefore, it can be seen that the dimension of each basic signal PCEN voiceprint feature is 128 dimensions.

[0199] From the above description, it can be seen that within a basic audio sliding window, the basic audio category vector feature is generated by the basic signal PCEN voiceprint feature. After the basic signal spectrogram is segmented by the basic audio sliding window, the corresponding basic audio category vector feature group is generated according to the PCEN voiceprint feature of each audio segment and the start and end time of the sound event marked in advance. Different from the basic signal PCEN voiceprint feature extracted above, the basic audio category vector feature group corresponding to each sound event can be obtained by foreground and background classification. In addition, the event category label y i It can be obtained by converting the above-mentioned annotations according to the category of the sound event.

[0200] For the above-mentioned frame-level labels, specifically: if a sound event lasts for 3s, with a start and end time of 1.12s to 4.12s respectively, and the time length corresponding to an audio frame is 25ms, then the frame indexes corresponding to the current sound event are 45 to 165. During the annotation conversion, the annotations for frames 45 to 165 are the categories of the event, and the rest are background categories.

[0201] According to the frame-level events, the small sample sound event detection basic model is learned on the classification branch task to prevent the foreground class c from being known in advance. t In one embodiment of the present invention, supervised training is performed on the first stage model of small sample sound event detection fine-tuning, and the characteristics of the basic audio sliding window can be used to select the foreground class c at the last moment. tThe basic audio sliding window of the target category forms an event time misalignment with the basic audio sliding window of the current moment; specifically, for a 20-second audio segment, when a window length of 5 seconds and a window shift of 1 second (corresponding to 86 audio frames) are performed, there will be 16 basic audio sliding windows. When processing the 8th basic audio sliding window, the corresponding foreground class, for example, class a, is used, and class a may also exist in the 7th basic audio sliding window or the 6th basic audio sliding window. If not, then the 8th basic audio sliding window is pending, the 9th basic audio sliding window is processed first, and then the target feature is extracted from the 8th basic audio sliding window, thereby forming a time misalignment with the 9th basic audio sliding window.

[0202] After the pre-training data set is constructed in the above manner, when the model is trained on the basic model of small sample sound event detection, generally, pre-training information should also be configured, wherein the configured pre-training information includes the target number of pre-training rounds and the pre-training loss function, that is, when the model is trained on the basic model of small sample sound event detection, the target number of pre-training rounds can be used as the termination condition of the pre-training. It should be understood that when all the pre-training samples in the pre-training data set are used to train the basic model of small sample sound event detection, one round of model training is completed.

[0203] After the pre-training of the basic model for small sample sound event detection reaches the pre-training target state, a small sample sound event detection pre-trained model can be generated based on the small sample sound event detection basic model. At this time, the corresponding network parameters of the feature extraction basic network, the information interaction basic network and the classifier basic network can be obtained through the small sample sound event detection pre-trained model.

[0204] Depend on Figure 3 It can be known that when pre-training the basic model for small sample sound event detection, for each loaded pre-training sample, the pre-training sample includes a basic signal PCEN voiceprint feature group and a basic audio category vector feature group. It should be noted that the basic signal PCEN voiceprint feature group and the basic audio category vector feature group in the same pre-training sample should be generated based on the same basic audio sliding window. During pre-training, the basic signal PCEN voiceprint feature group and the basic audio category vector feature group of the pre-training sample are loaded in parallel into the feature extraction basic network. Thereafter, the basic signal PCEN voiceprint feature group and the basic audio category vector feature group can be respectively extracted through the feature extraction basic network.

[0205] From the above description, it can be seen that based on the same basic audio sliding window, a basic signal PCEN voiceprint feature group and multiple basic audio category vector feature groups can be obtained. Then a basic audio category vector feature group and a basic signal PCEN voiceprint feature group constitute a pre-training sample, that is, based on the same basic audio sliding window, multiple pre-training samples can be formed, which is related to the number of different sound events corresponding to the basic audio sliding window. Please refer to the above description for details.

[0206] In specific implementation, the pre-training loss function may adopt the cross entropy loss. When the pre-training loss function adopts the cross entropy loss, then:

[0207]

[0208] Among them, X i is the basic signal PCEN voiceprint feature group in the i-th pre-training sample, Y i =(y1,y2,…,y n ) is the frame-level label corresponding to the basic audio category vector feature in the i-th pre-training sample, y i ∈{0,…,m}, “1~m” indicates that there are m types of sound event labels in the pre-training dataset, “0” is the background event class, and b is the total number of event categories in the pre-training dataset (b=m+1); X ith is the basic audio category vector feature group in the i-th pre-training sample, A i =(a1,a2,…,a i ), a i ∈{0,1} is the foreground / background label corresponding to the basic audio category vector feature, M i ∈{0,1} is the training mask.

[0209] Specifically, Θ represents the dot product operation, f φ is the basic model for small sample sound event detection, f φ (X i ) is the predicted output of the basic model for small sample sound event detection for the frame-level label corresponding to the basic audio category vector feature in the i-th pre-training sample, N is the number of pre-training samples; n is the number of audio frames corresponding to a basic audio sliding window.

[0210] For the above pre-training loss function, initially, Y1==1, which means all frame-level labels with sound event label 1 in the first pre-training sample; if Y2=1, it means all frame-level labels with sound event label 1 in the second pre-training sample; A1==1, which means all frame-level labels belonging to the foreground class in the first pre-training sample. φ (X i ,X ith) represents the predicted output of the foreground and background frame level labels corresponding to the basic audio category vector features in the i-th pre-training sample.

[0211] In the specific implementation, the pre-training loss function adopts the cross entropy loss, and the network parameters of the basic model of small sample sound event detection can be updated by the stochastic gradient descent method. Furthermore, for two adjacent basic audio sliding windows, there will be overlaps, and the training mask M is used. i The overlapping part of the same event in the adjacent basic audio sliding windows is configured not to calculate the loss, where the training mask M i When a sound event is captured by two adjacent basic audio sliding windows at the same time, the training mask of the part of the overlapping event corresponding to the latter basic audio sliding window and the former basic audio sliding window is 0, otherwise it is 1.

[0212] In one embodiment of the present invention, the first-stage fine-tuning model for small sample sound event detection includes a first-stage feature extraction network, a first-stage information interaction network, and a first-stage classifier network, wherein:

[0213] The first-stage feature extraction network is generated by migrating the feature extraction basic network in the pre-trained model of small-sample sound event detection;

[0214] The first stage network of information interaction is generated by migrating the basic network of information interaction in the model pre-trained by small sample sound event detection;

[0215] The first-stage classifier network includes a first-stage first classifier, a first-stage second classifier, and a first-stage third classifier, wherein:

[0216] In the first stage, the first classifier is generated by transferring the base class basic classifier in the pre-trained model of small sample sound event detection;

[0217] The second classifier in the first stage is generated by transferring the basic foreground and background classifiers in the pre-trained model of small sample sound event detection;

[0218] The third classifier in the first stage is constructed based on the prototype feature extraction method;

[0219] The first classifier of the first stage is connected with the network of the first stage of feature extraction through adaptation;

[0220] The first-stage second classifier, the first-stage third classifier and the first-stage information interaction network adapter are connected.

[0221] Specifically, after obtaining the pre-trained model for small sample sound event detection, the network parameters of the basic network for feature extraction can be obtained according to the network parameters of the first-stage network for feature extraction, that is, the network parameters of the first-stage network for feature extraction are consistent with the network parameters of the basic network for feature extraction. Similarly, the network parameters of the first-stage network for information interaction are consistent with the corresponding network parameters of the basic network for signal interaction, that is, the migration generation of the present invention is realized. In addition, the corresponding migration generation meanings of the first-stage first classifier and the first-stage second classifier can also be obtained.

[0222] From the above description, it can be seen that the third classifier in the first stage is a new classifier constructed by fine-tuning the first stage model in small sample sound event detection. Figure 8 An embodiment of the first stage model of small sample sound event detection fine-tuning is shown in the figure, wherein stage one refers to the first stage model of small sample sound event detection fine-tuning, 4×CNN corresponds to the first stage network for feature extraction, the information interaction network corresponds to the first stage network for information interaction, Decoder3 is the first stage first classifier, Decoder4 is the first stage second classifier, and Decoder5 is the first stage third classifier.

[0223] In one embodiment of the present invention, the small sample fine-tuning stage training data set includes a fine-tuning first stage training data set and a fine-tuning second stage training data set, wherein:

[0224] The fine-tuning first stage model for small sample sound event detection is trained using the fine-tuning first stage training dataset;

[0225] The fine-tuning second stage model for small sample sound event detection is trained using the fine-tuning second stage training dataset;

[0226] The first stage of fine-tuning training training datasets include mixed domain datasets, single sound event support datasets, and multiple sound event support datasets;

[0227] When constructing the first stage training dataset for fine-tuning training, it includes:

[0228] Providing a registered audio set matching the target audio signal, wherein the registered audio set includes a plurality of registered audios, and each registered audio is a strongly marked audio;

[0229] Performing a third audio signal preprocessing on each registered audio to generate a registered audio feature set, wherein the registered audio feature set includes a plurality of registered audio PCEN voiceprint feature groups with time series features and a registered audio category vector feature group corresponding to each registered audio PCEN voiceprint feature,

[0230] Based on a registered audio PCEN voiceprint feature group and a corresponding registered audio category vector feature group, a registered audio training sample is formed;

[0231] Sampling a target number of pre-training samples in the pre-training dataset, and mixing the sampled pre-training samples with all registered audio training samples to form a mixed domain dataset;

[0232] Based on the registered audio feature set, a single sound event support dataset and multiple sound event support datasets are generated;

[0233] When performing model training on the small sample sound event detection fine-tuning first stage model, the corresponding first stage training samples in the mixed domain dataset, the single sound event support dataset, and the multiple sound event support datasets are respectively loaded into the small sample sound event detection fine-tuning first stage model to perform model training on the small sample sound event detection fine-tuning first stage model.

[0234] It should be noted that the registered audio set matches the target audio signal, which specifically means that each registered audio in the registered audio set and the target audio signal belong to the audio signal in the same scene, but each registered audio belongs to the strongly labeled audio. It can be understood that the above-mentioned small sample specifically refers to the small number of registered audios in the registered audio set. Generally, there is only one registered audio in the registered audio set, and the number of sound events contained in the registered audio is small, such as there may be only 5 registered sound events in one registered audio. Of course, the number of registered audios and registered sound events in the registered audio can be collected and determined according to the actual scene.

[0235] For the registered audio in the registered audio set, perform the third preprocessing of the audio signal, wherein after performing the third preprocessing of the audio signal, a registered audio feature set can be generated, and the registered audio feature set includes a registered audio PCEN voiceprint feature group and a corresponding registered audio category vector feature group; the manner and process of performing the third preprocessing of the audio signal can refer to the above-mentioned description of performing the second preprocessing of the audio signal, which will not be repeated here. For the registered audio feature set, based on a registered audio PCEN voiceprint feature group and a corresponding registered audio category vector feature group, a registered audio training sample is formed. It should be noted that the registered audio PCEN voiceprint feature group and the registered audio category vector feature group in the registered audio training sample should be formed based on the same segmentation window. The segmentation window can refer to the corresponding description of the above-mentioned target signal sliding window and basic signal sliding window, which will not be repeated here. It can be seen that based on the registered audio feature set, multiple registered audio training samples can be obtained.

[0236] A target number of pre-training samples are sampled in the pre-training data set, and the sampled pre-training samples are mixed with all the registered audio training samples to form a mixed domain data set, that is, the mixed data set includes both the registered audio training samples and the pre-training samples. In specific implementation, the target number of pre-training samples can be formed by randomly sampling 10% of the pre-training samples from the pre-training data set. Of course, the target number of sampled pre-training samples can also be other situations, which can be selected according to needs. It can be seen that the mixed domain data set includes a number of registered audio training samples and a number of pre-training samples.

[0237] From the above description, it can be seen that since the number of sound events in the registered audio is small, the obtained registered audio training samples are also small. In order to meet the needs of fine-tuning the first stage model of small sample sound event detection for model training, the training sample expansion operation should be performed. In specific implementation, the registered audio samples can be expanded by event resampling. The event resampling specifically refers to repeated sampling of the registered sound events in the registered audio set to simulate real scenes and expand the number of training samples.

[0238] During the training sample expansion operation, first, on the registered audio spectrogram of the registered audio, according to the annotation information (the annotation information is the category of the sound event in the registered audio and the start and end time of the sound event), the registered event audio (pos) and the non-registered event audio (neg) mixed in the registered event audio are extracted, and the non-registered event audio neg segment with a longer duration is segmented using a sliding window to obtain a negative class sample set (negs). It should be noted that in the audio recorded based on the real scene, there will be a long period of time when it is background non-target sound. The non-target sound is the non-registered event audio neg segment mentioned above, and the registered event audio is the sound event determined according to the annotation. By performing sliding window segmentation on the non-registered event audio neg segment, multiple negative class samples can be obtained, and thus a negative class sample set can be formed, wherein, when the sliding window segmentation is performed, it is necessary to construct a filling sample window, and the window length of the filling sample window should be consistent with the corresponding window length of the target signal sliding window and the basic signal sliding window.

[0239] For a registered audio, after obtaining the negative sample set, at least one registered event audio is randomly selected from all the registered event audios, and the selected registered event audio is randomly filled in a filling sample window sampled by a non-registered event audio negs, and a filling audio signal is formed, and the PCEN voiceprint features corresponding to the registered event audio are filled in the filling audio signal, so as to simulate the scenario of a single sound event and multiple sound events in the same scene; then, the above steps are repeated multiple times to form a single sound event support data set and multiple sound event support data set representations respectively, Figure 5An embodiment of a single sound event support data set is shown in Figure 6 An embodiment of multiple sound event support data sets is shown in FIG. It should be noted that the length of the PCEN voiceprint feature group corresponding to each registration event audio is less than the window length of the filling sample window.

[0240] In specific implementation, a single sound event specifically refers to using only one registered sound event when performing random area filling; multiple sound events specifically refer to using multiple registered sound events at the same time when performing random area filling. It should be noted that the formed single sound event support data set and multiple sound event support data set contain the corresponding registered audio category vector features; specifically, the registered sound event specifically refers to the sound event in the registered audio.

[0241] The formed single sound event support data set includes several single sound event support data, wherein a single sound event, such as a call, is a strongly labeled audio, therefore, the category of the single sound event and the start and end positions of the sound event are known, a single sound event support data forms a single sound event support sample, a single sound event support sample is a single sound event PCEN voiceprint feature group, a single sound event PCEN voiceprint feature group includes several single sound event PCEN voiceprint features, as can be seen from the above description, it can include 431 single sound event PCEN voiceprint features. For the case of multiple sound event support data sets, please refer to the description of the single sound event support data set here, which will not be repeated here.

[0242] In order to prevent the registration event audio pos in the filling sample window from lasting too short, resulting in serious imbalance between the registration event audio pos and the non-registration event audio neg in the filling audio signal, during the filling process of multiple registration event audio pos, the registration sound event is repeatedly sampled to ensure that the length of the registration event audio pos in a filling audio signal formed after filling accounts for at least 67% of the total length of the filling sample window.

[0243] Considering that the event resampling method cannot change the audio information during the occurrence of the sound event, and multiple resampling of the same registered event audio is prone to overfitting, the data enhancement method of the linear filter with a temporal mask is used to enhance the single sound event support data and multiple sound event support data, such as Figure 7 The specific steps are as follows:

[0244] First, the filling sample window (as can be seen from the above description, the filling sample window can correspond to 431 audio frames) is randomly divided into m parts (in one embodiment of the present invention, m is set to a random natural number between 3 and 6), denoted as T = {T1, T2, ..., Tm}, Figure 7 An embodiment of a linear filter when m is 4 is shown in FIG. 1 ; Next, for each T i Assign a random gain factor g ranging from 0 to 1 i ; Then, according to each g i The corresponding T i The linear enhancement of -6dB to 8dB is performed. The linear enhancement of the linear filter can be performed in the following manner:

[0245]

[0246] Where, “·” represents element-wise multiplication, β is the dB gain factor, and l x With r s are the gain lower bound and the gain upper bound of the linear filter, respectively. In one embodiment of the present invention, the gain lower bound is l x Can be set to -6, gain upper boundary r s Can be set to 8, γ represents two adjacent gain factors g i The linear value interval between the two is the length of the interval and the corresponding T i The interval lengths are consistent.

[0247] It should be noted that both single sound event support data and multiple sound event support data should be linearly enhanced by the above-mentioned linear filter. During linear enhancement, the amplitude of the corresponding PCEN voiceprint feature in each filled sample window can be enhanced.

[0248] In one embodiment of the present invention, the third classifier of the first stage can be constructed based on the prototype feature extraction method, wherein the method of using a single sound event support data set and constructing the third classifier of the first stage based on the prototype feature extraction method specifically includes:

[0249]

[0250] Among them, W p is the initial weight matrix of the third classifier in the first stage, nn is the total number of sliding windows when the registered audio is segmented by sliding windows, and λ i is the PCEN voiceprint feature group of the i-th registered audio in the single sound event support dataset, f φ ′ is the basic network for feature extraction.

[0251] Specifically, f φ ′(λ i ) represents the use of feature extraction network to extract the PCEN voiceprint feature set λ of the registered audio i The feature extraction information generated by feature extraction, I(χ i =0), I(χ i=1) is an exponential function, χ i =0 means that the audio corresponding to the i-th segmentation sliding window is background sound. i =0, I(χ i =0) = 1, that is, the audio corresponding to the i-th segmentation sliding window belongs to the background, otherwise it is 0; I(χ i =1) can refer to I(χ i =0) corresponding instructions.

[0252] The initial weight matrix W of the third classifier in the first stage p within The molecular part is the background χ for all labels i = 0 using the feature extraction network to extract the PCEN voiceprint feature set λ of the registered audio i Extract features and accumulate the features of all background windows. The denominator counts the total number of background windows, and the result represents the feature prototype of the background class window, that is, the average value of the feature vectors of all background windows.

[0253] From the above description, it can be seen that the first-stage training dataset constructed for fine-tuning training includes a mixed domain dataset, a single sound event support dataset, and a multiple sound event support dataset; Figure 8 In the example, the support set includes a single sound event support dataset and multiple sound event support datasets; thereafter, when training the first-stage model of small-sample sound event detection fine-tuning, the corresponding data samples in the mixed domain dataset, the single sound event support dataset, and the multiple sound event support dataset are sequentially sent to the first-stage model of small-sample sound event detection fine-tuning in three ways to obtain three-way output features, which are: X t , X s and X p ; For example, based on the mixed domain dataset, after feature extraction in the first stage of feature extraction network, the output feature X can be obtained t ; For a single sound event support data set, after feature extraction in the first stage network, the output feature X can be obtained s , after extracting features from the first-stage network for multiple sound event support datasets, the output feature X can be obtained p .

[0254] When we get the output feature X t , output feature X s , output feature X p Afterwards, they are sent into the first stage network of information interaction respectively, and the model is fine-tuned with different task branches. It can be understood that the model fine-tuning here is supervised training.

[0255] Specifically, the output feature Xt Mainly used for fine-tuning of mixed domain models, output feature X s With the output feature X p They are used for optimization of the foreground and background classification branch and the event detection task branch during back propagation. Before the final output, the first stage network of capturing information interaction is used to realize the information interaction between the two branches and the capture of context information. Among them, the output feature X p The input data loaded into the Query end of the first phase of information interaction is output as feature X s Input data to the Key and Value ends of the first phase of information interaction network; From the above description, we can see that the output feature X s It should be transferred through the SED branch vector module and loaded into the Key and Value ends of the first stage network of information interaction, and the output feature X p The FBSC branch vector module should be used, and the above-mentioned repeat recovery operation and POS prototype module should be loaded into the Query end of the first stage network of information interaction.

[0256] The hybrid domain fine-tuning strategy is: the output feature Xt outputs the corresponding first prediction label through the first stage first classifier, and then the cross entropy loss is calculated using the real hybrid domain label, wherein the calculation of the cross entropy loss can refer to the calculation description of l1 in the above pre-training loss. The output feature Xs is loaded into the first stage network of information interaction, and then the corresponding second prediction label is output by the first stage second classifier, and the second prediction label is cross-entropy loss calculated with the foreground and background labels, wherein the calculation of the cross entropy loss can refer to the calculation description of l2 in the above pre-training loss. Similarly, the output feature Xp is loaded into the first stage network of information interaction, and then the corresponding third prediction label is output by the first stage third classifier, and the third prediction label is cross-entropy loss calculated with the event label, wherein the calculation of the cross entropy loss can refer to the calculation description of l2 in the above pre-training loss. In specific implementation, after the above three cross entropy losses are calculated, the three cross entropy losses are correspondingly accumulated, thereby obtaining the total loss of the first stage of fine-tuning.

[0257] After determining the total loss of the first stage of fine-tuning, the model parameters of the first stage model of small sample sound event detection fine-tuning are updated by the chain rule. The method of updating the model parameters can be consistent with the prior art. For example, the corresponding gradient weights of each first stage first classifier to first stage third classifier can be calculated by back propagation. Thereafter, the calculated gradient weights are used to update the weights of the first stage model of small sample sound event detection fine-tuning through the optimizer, so that the loss value of the first stage model of small sample sound event detection fine-tuning is reduced during the next forward propagation. In order to enhance the feature extraction and classification capabilities, the process of updating the model parameters of the first stage model of small sample sound event detection fine-tuning according to the chain rule can be selected according to actual needs and will not be repeated here.

[0258] It should be noted that after mixed domain fine-tuning, the small sample sound event detection fine-tuning second stage model is trained. Specifically, the small sample sound event detection fine-tuning second stage model can be obtained through weight sharing. The weight update can refer to the corresponding migration generation instructions in the following small sample sound event detection fine-tuning second stage model, which will not be repeated here.

[0259] In one embodiment of the present invention, the small sample sound event detection fine-tuning second stage model includes a feature extraction second stage network, an information interaction second stage network and a classifier second stage network, wherein:

[0260] The second-stage network for feature extraction is generated by migrating the basic network for feature extraction in the pre-trained model for small-sample sound event detection.

[0261] The second-stage network of information interaction is generated by migrating the basic network of information interaction in the model pre-trained for small sample sound event detection;

[0262] The second stage network of the classifier includes a second stage first classifier and a second stage second classifier, wherein:

[0263] The first classifier in the second stage is generated by migrating the second classifier in the first stage;

[0264] The second classifier in the second stage is generated by migrating the third classifier in the first stage;

[0265] The first classifier of the second stage, the second classifier of the second stage and the second network adaptation connection of the information interaction stage;

[0266] When constructing the second stage training dataset for fine-tuning, it includes:

[0267] Acquire a reference audio signal matching the target audio signal, wherein the reference audio signal is a weakly labeled audio signal;

[0268] Performing a first audio signal preprocessing on the reference audio to generate a reference signal feature set of the reference audio signal, wherein the reference signal feature set includes a plurality of reference signal PCEN voiceprint feature groups with time series characteristics, and each reference signal PCEN voiceprint feature group is used as a fine-tuning second-stage training sample in a fine-tuning second-stage training data set;

[0269] When training the fine-tuning second-stage model for small-sample sound event detection, the time series characteristics are used to load the fine-tuning second-stage training samples one by one into the feature extraction second-stage network to extract reference signal voiceprint extraction features;

[0270] In the second stage of information interaction, the network interacts with the reference signal voiceprint extraction feature group and the multi-sound event training feature group, where:

[0271] The multi-sound event training feature group is generated by extracting and migrating the features of the first-stage training samples corresponding to the multiple sound event support data sets using the first-stage feature extraction network in the current epoch fine-tuning training;

[0272] The reference signal voiceprint extraction features are loaded into the frame-level first input channel of the second-stage network of information interaction;

[0273] The multi-sound event training features are loaded into the frame-level second input channel of the second-stage network of information interaction;

[0274] The information interaction first output information and the information interaction second output information of the information interaction second stage network are fused and output by the classifier second stage network to generate a predicted frame probability for each audio frame in the current fine-tuning second stage training sample;

[0275] Using the adaptive pseudo-label threshold thresh and the predicted frame probability of each audio frame, the fine-tuning second-stage training samples are configured as high-confidence pseudo-labels or low-confidence pseudo-labels, and the fine-tuning second-stage training samples corresponding to the high-confidence pseudo-labels are added to the fine-tuning training first-stage training data set to update the fine-tuning training first-stage training data set.

[0276] It should be noted that the reference audio signal matches the target audio signal, which specifically means that the reference audio signal can be the target audio signal, or a weakly labeled audio in the same scene as the target audio signal; in specific implementation, the reference audio signal can preferably be a segment of the target audio signal, that is, the reference audio signal belongs to the target audio signal. Of course, when the target audio signal is used as the reference audio signal, the trained small sample sound time detection model will not affect the subsequent sound event detection of other target audio signals.

[0277] Figure 8Also shown is an embodiment of the second-stage model for fine-tuning small sample sound event detection. In the figure, stage two is the second-stage model for fine-tuning small sample sound event detection, the 4×CNN in stage two is the second-stage network for feature extraction, the information interaction network is the second-stage network for information interaction, Decoder6 is the first classifier of the second stage, and Decoder7 is the first classifier of the second stage.

[0278] In the second stage of training, the unlabeled reference signal PCEN voiceprint feature group is sent to the feature extraction second stage network for feature extraction to obtain feature X q ; Then, the output features X extracted from the samples in the first stage of the multiple sound event support data set are p Transfer to the second stage and get feature X ps ; Then, feature X ps With feature X q are sent together to the second stage network of information interaction, where feature X ps Loaded to the query side of the second stage network of information interaction, feature X q Loaded into the Value and KEY ends of the second phase of information interaction network.

[0279] In the specific implementation, transfer to the second stage to obtain feature X ps When all the output features X p Perform the mean operation to generate feature X after the mean operation ps Since the first stage training and the second stage training are performed alternately in one round of training, all the output features X obtained by the first stage training are p After taking the mean operation, feature X can be generated ps ; In the next round of training, all output features X are used in the same way p After taking the mean operation, feature X can be generated ps In actual reasoning, the feature X obtained in the last round of training can be ps As a reference feature for multiple sound events.

[0280] Since the output space of the event classification branch is consistent with that of the foreground and background classification branch, the output probability results of the two branches are averaged and fused by output fusion; finally, the high / low confidence pseudo-labels are distinguished according to the adaptive pseudo-label threshold thresh, and the high-confidence pseudo-labels are added to the first stage training of the next round, that is, sent to the single sound event support data set for the first stage training of the next round. Specifically, the output probability is compared with the adaptive pseudo-label threshold thresh to label the predicted data (pseudo-label). Therefore, each predicted data will have a pseudo-label, but not all data are used. It is necessary to reuse the probability value output by the network and take the pseudo-label with a relatively high output probability for use. In one embodiment of the present invention, the screening of "relatively high output probability" is to sort all predicted probabilities and take the top 20% sample windows.

[0281] During the fine-tuning process, the two stages are trained alternately according to the set number of training rounds. For example, in one round of training, the first stage of training is performed first, and then the second stage of training. When the reference audio signal belongs to the target audio signal, in the final prediction, the output result of the second stage in the last round can be used as the final prediction, and the predicted frame with an output probability exceeding 0.5 is regarded as "1", otherwise it is "0".

[0282] In one embodiment of the present invention, the two-stage fine-tuning strategy is as follows: during the first stage of training, the third classifier of the first stage is supervisedly trained using the registered audio with data enhancement, and then the trained classifier is used to pseudo-label subsequent audio events. During the second stage of training, the pseudo-labels output by the first stage are sent to the detection network for semi-supervised training, and the two stages are alternately learned until the training termination condition is met.

[0283] In each round of fine-tuning, the pseudo-label data is adaptively annotated according to the recognition ability of the target sound event, and its mathematical expression is:

[0284]

[0285] Among them, P is the sample softmax output probability value, v is the distance hyperparameter, thre is the adaptive threshold, and the distance hyperparameter v can generally be taken as 0.3.

[0286] In one embodiment of the present invention, the adaptive threshold may be calculated as follows:

[0287]

[0288] Where j is the window index, N s Represents the number of single sound event support data in a single sound event support dataset, n j is the window length of the filling sample window, xi Supports data for a single sound event, h i It is a pseudo label corresponding to each PCEN voiceprint feature in the single sound event support data.

[0289] From the above description, we can see that the window length n of the filled sample window j Generally, it can be 431 frames, and a single sound event supports data x i The window length n may include the padded sample window j Consistent PCEN voiceprint features, such as window length n j When the frame is 431, a single sound event supports data x i It should contain 431 corresponding PCEN voiceprint features.

Claims

1. A small sample sound event detection method based on frame level, characterized in that: The small sample sound event detection method comprises: Providing a target audio signal to be detected, and performing a first audio signal preprocessing on the target audio signal to extract a target signal feature set of the target audio signal, wherein the target signal feature set includes a plurality of target signal PCEN voiceprint feature groups with time series characteristics, the target signal PCEN voiceprint feature groups include a plurality of target signal PCEN voiceprint features, and one target signal PCEN voiceprint feature corresponds to one audio frame of the target audio signal; The extracted target signal feature set is loaded into the constructed small sample sound event detection model, so as to use the small sample sound event detection model to perform sound event detection on the target audio signal, wherein: When performing sound event detection on the target audio signal, the target signal PCEN voiceprint feature groups are loaded into the small sample sound event detection model one by one according to the time series characteristics; Performing detection processing on each target signal PCEN voiceprint feature group using a small sample sound event detection model, so as to generate a predicted frame state corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group after the detection processing; Based on the predicted frame states corresponding to the PCEN voiceprint features of all target signals, determine the sound event categories contained in the target audio signal and the time start and end points corresponding to each sound event category; For each target signal PCEN voiceprint feature, we have: Wherein, PCEN(t,f) is the target signal PCEN voiceprint feature, t is the time of the audio frame in the target audio signal corresponding to the current target signal PCEN voiceprint feature, f is the Mel spectrum response corresponding to the current audio frame in the target audio signal, E(t,f) is the input spectrum energy mean corresponding to the current audio frame in the target audio signal, ξ is a decimal constant factor, α is a gain factor, α∈(0,1), δ and γ are output range compression factors, and M(t,f) is a first-order IIR filter; The small sample sound event detection model includes a feature extraction network, an information interaction network, a classifier network and a prediction frame splicing module connected in sequence, wherein: When detecting and processing each target signal PCEN voiceprint feature group, a feature extraction network is used to extract features of the target signal PCEN voiceprint feature group to obtain target signal voiceprint extraction features of the target signal PCEN voiceprint feature group; The information interaction network interacts with the target signal voiceprint extraction features and the multiple sound event reference features to generate information interaction first output information and information interaction second output information; The classifier network fuses the first output information of the information interaction and the second output information of the information interaction, and determines the prediction frame state corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group through the prediction frame splicing module; For all target signal PCEN voiceprint feature groups, based on the predicted frame state corresponding to each target signal PCEN voiceprint feature in each target signal PCEN voiceprint feature group, the predicted frame splicing module determines all sound event categories of the target audio signal and the time start and end points corresponding to each sound event category.

2. The frame-level small sample sound event detection method according to claim 1 is characterized in that: The information interaction network includes a frame-level input unit, a multi-head self-attention unit, and a secondary feature fusion unit, wherein: The frame-level input unit includes a frame-level first input channel and a frame-level second input channel, wherein the target signal voiceprint extraction feature is loaded into the frame-level first input channel, and the multi-sound event reference feature is loaded into the frame-level second input channel; Based on the target signal voiceprint extraction feature, the frame-level first input channel generates frame-level first channel output information, wherein the frame-level first channel output information directly jumps to output and forms the information interaction first output information of the information interaction network, and the frame-level first channel output information is also loaded to the Key end and the Value end of the multi-head self-attention unit; Based on the multi-sound event reference features, the frame-level second input channel generates frame-level second channel output information, and the frame-level second channel output information is also loaded into the Query end of the multi-head self-attention unit; The multi-head self-attention unit performs initial attention calculation on the frame-level first channel output information loaded via the Key end, the frame-level first channel output information loaded via the Value end, and the frame-level second channel output information loaded via the Query end, so as to generate attention feature information after the initial attention calculation; Additively fusing the attention feature information with the frame-level first channel output information generated by the frame-level first input channel to form information interaction preliminary fusion information, and loading the information interaction preliminary fusion information into the secondary feature fusion unit; The information interaction preliminary fusion information and the secondary feature fusion unit adopt residual connection to form information interaction secondary fusion information, and the information interaction secondary fusion information is used as the second output information of information interaction.

3. The frame-level small sample sound event detection method according to claim 2 is characterized in that: The classifier network includes a foreground and background classifier, an event classifier, and a classifier splicer, wherein: The foreground and background classifier receives the first output information of the information interaction, and the event classifier receives the second output information of the information interaction; The foreground and background classifiers, the event classifiers and the classification splicer are adapted and connected; The first output information of the information interaction and the second output information of the information interaction are fused through the foreground and background classifier, the event classifier and the classification splicer to generate the prediction probability corresponding to each target signal PCEN voiceprint feature in the current target signal PCEN voiceprint feature group; For each prediction probability, the prediction frame splicing module generates a prediction frame state corresponding to each audio frame based on a preset prediction probability threshold.

4. The frame-level small sample sound event detection method according to claim 3 is characterized in that: When building a small sample sound event detection model, include: Build a basic model for small sample sound event detection, Constructing a pre-training data set, and using the pre-training data set to pre-train a basic model for small sample sound event detection, so as to generate a small sample sound event detection pre-trained model when the pre-training reaches a pre-training target state; Based on the pre-trained model of small sample sound event detection, a first-stage fine-tuning model of small sample sound event detection and a second-stage fine-tuning model of small sample sound event detection are constructed; A small sample fine-tuning stage training data set is constructed, and the small sample fine-tuning stage training data set is used to fine-tune the small sample sound event detection fine-tuning first stage model and the small sample sound event detection fine-tuning second stage model, wherein: When performing fine-tuning training for each epoch, supervised training is first performed on the first stage model of small sample sound event detection fine-tuning, and then semi-supervised training is performed on the second stage model of small sample sound event detection fine-tuning; When the fine-tuning training reaches the fine-tuning training target state, the small sample sound event detection fine-tuning second stage model configuration that reaches the fine-tuning training target state is used as the small sample sound event detection model.

5. The frame-level small sample sound event detection method according to claim 4 is characterized in that: The small sample sound event detection basic model includes a feature extraction basic network, an information interaction basic network and a classifier basic network, wherein: The feature extraction basic network, the information interaction basic network and the classifier basic network are connected in sequence, and the classifier basic network includes a base class basic classifier and a foreground and background basic classifier; When building a pre-training dataset, include: Produce a large-scale audio data set, wherein the large-scale audio data set includes a plurality of basic audio signals, each of which is a strongly labeled audio; Performing a second audio signal preprocessing on each basic audio signal to generate a basic audio feature set, wherein the basic audio feature set includes a plurality of basic signal PCEN voiceprint feature groups with time series characteristics and a basic audio category vector feature group corresponding to each basic signal PCEN voiceprint feature group; A pre-training data set is constructed based on a basic audio feature set of a basic audio signal, wherein for any pre-training sample in the pre-training data set, it includes a basic signal PCEN voiceprint feature group and a basic audio category vector feature group corresponding to the basic signal PCEN voiceprint feature group; Configure pre-training information and use the pre-training data set to train a basic model for small sample sound event detection, where: The configured pre-training information includes the target number of pre-training rounds and the pre-training loss function; When the number of training rounds of the small sample sound event detection basic model using the pre-training data set reaches the pre-training target number of rounds, the pre-training of the small sample sound event detection basic model reaches the pre-training target state, and a small sample sound event detection pre-trained model is generated based on the small sample sound event detection basic model.

6. The frame-level small sample sound event detection method according to claim 5 is characterized in that: The first-stage fine-tuning model for small sample sound event detection includes a first-stage feature extraction network, a first-stage information interaction network, and a first-stage classifier network, wherein: The first-stage feature extraction network is generated by migrating the feature extraction basic network in the pre-trained model of small-sample sound event detection; The first stage network of information interaction is generated by migrating the basic network of information interaction in the model pre-trained by small sample sound event detection; The first-stage classifier network includes a first-stage first classifier, a first-stage second classifier, and a first-stage third classifier, wherein: In the first stage, the first classifier is generated by transferring the base class basic classifier in the pre-trained model of small sample sound event detection; The second classifier in the first stage is generated by transferring the basic foreground and background classifiers in the pre-trained model of small sample sound event detection; The third classifier in the first stage is constructed based on the prototype feature extraction method; The first classifier of the first stage is connected with the network of the first stage of feature extraction through adaptation; The first-stage second classifier, the first-stage third classifier and the first-stage information interaction network adapter are connected.

7. The frame-level small sample sound event detection method according to claim 6 is characterized in that: The small sample fine-tuning stage training data set includes a fine-tuning first stage training data set and a fine-tuning second stage training data set, wherein: The fine-tuning first stage model for small sample sound event detection is trained using the fine-tuning first stage training dataset; The fine-tuning second stage model for small sample sound event detection is trained using the fine-tuning second stage training dataset; The first stage of fine-tuning training training datasets include mixed domain datasets, single sound event support datasets, and multiple sound event support datasets; When constructing the first stage training dataset for fine-tuning training, it includes: Providing a registered audio set matching the target audio signal, wherein the registered audio set includes a plurality of registered audios, and each registered audio is a strongly marked audio; Performing a third audio signal preprocessing on each registered audio to generate a registered audio feature set, wherein the registered audio feature set includes a plurality of registered audio PCEN voiceprint feature groups with time series features and a registered audio category vector feature group corresponding to each registered audio PCEN voiceprint feature, Based on a registered audio PCEN voiceprint feature group and a corresponding registered audio category vector feature group, a registered audio training sample is formed; Sampling a target number of pre-training samples in the pre-training dataset, and mixing the sampled pre-training samples with all registered audio training samples to form a mixed domain dataset; Based on the registered audio feature set, a single sound event support dataset and multiple sound event support datasets are generated; When performing model training on the small sample sound event detection fine-tuning first stage model, the corresponding first stage training samples in the mixed domain dataset, the single sound event support dataset, and the multiple sound event support datasets are respectively loaded into the small sample sound event detection fine-tuning first stage model to perform model training on the small sample sound event detection fine-tuning first stage model.

8. The frame-level small sample sound event detection method according to claim 7 is characterized in that: The small sample sound event detection fine-tuning second stage model includes a feature extraction second stage network, an information interaction second stage network and a classifier second stage network, wherein: The second-stage network for feature extraction is generated by migrating the basic network for feature extraction in the pre-trained model for small-sample sound event detection. The second-stage network of information interaction is generated by migrating the basic network of information interaction in the model pre-trained for small sample sound event detection; The second stage network of the classifier includes a second stage first classifier and a second stage second classifier, wherein: The first classifier in the second stage is generated by migrating the second classifier in the first stage; The second classifier in the second stage is generated by migrating the third classifier in the first stage; The first classifier of the second stage, the second classifier of the second stage and the second network adaptation connection of the information interaction stage; When constructing the second stage training dataset for fine-tuning, it includes: Acquire a reference audio signal matching the target audio signal, wherein the reference audio signal is a weakly labeled audio signal; Performing a first audio signal preprocessing on the reference audio to generate a reference signal feature set of the reference audio signal, wherein the reference signal feature set includes a plurality of reference signal PCEN voiceprint feature groups with time series characteristics, and each reference signal PCEN voiceprint feature group is used as a fine-tuning second-stage training sample in a fine-tuning second-stage training data set; When training the fine-tuning second-stage model for small-sample sound event detection, the time series characteristics are used to load the fine-tuning second-stage training samples one by one into the feature extraction second-stage network to extract reference signal voiceprint extraction features; In the second stage of information interaction, the network interacts with the reference signal voiceprint extraction feature group and the multi-sound event training feature group, where: The multi-sound event training feature group is generated by extracting and migrating the features of the first-stage training samples corresponding to the multiple sound event support data sets using the first-stage feature extraction network in the current epoch fine-tuning training; The reference signal voiceprint extraction features are loaded into the frame-level first input channel of the second-stage network of information interaction; The multi-sound event training features are loaded into the frame-level second input channel of the second-stage network of information interaction; The information interaction first output information and the information interaction second output information of the information interaction second stage network are fused and output by the classifier second stage network to generate a predicted frame probability for each audio frame in the current fine-tuning second stage training sample; Using the adaptive pseudo-label threshold thresh and the predicted frame probability of each audio frame, the fine-tuning second-stage training samples are configured as high-confidence pseudo-labels or low-confidence pseudo-labels, and the fine-tuning second-stage training samples corresponding to the high-confidence pseudo-labels are added to the fine-tuning training first-stage training data set to update the fine-tuning training first-stage training data set.