Audio Detection Model Training and Audio Detection Method, Device, Equipment and Medium
Through the multimodal combined audio detection model, multimodal feature extraction and model training of audio sample data is solved, and the problem of low accuracy of existing audio detection methods is achieved, and higher audio detection accuracy is achieved.
Patent Information
- Application Number
- CN202111209534.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-10-18
AI Technical Summary
Existing audio detection methods only perform audio detection based on single sample type or single feature, resulting in low detection accuracy.
By acquiring audio sample data including at least two audio sample subdata, multimodal audio features are extracted using a multimodal joint audio detection model, and a multimodal joint audio detection model is trained for audio detection.
It improves the accuracy of audio detection and solves the problem of low accuracy under single sample type or single feature detection.
Smart Images

Figure CN113903359B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and further relates to artificial intelligence fields such as deep learning and audio processing. Specifically, it relates to an audio detection model training and audio detection method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid improvement of current artificial intelligence technology and computer hardware performance in recent years, breakthroughs have been made in application fields such as computer audition, natural language processing, and audio detection. As a basic task in the field of computer audition, the accuracy of audio detection has also been greatly improved and is becoming more and more widely used in all walks of life. Summary of the Invention
[0003] Embodiments of the present disclosure provide an audio detection model training and audio detection method, device, equipment, and medium, which can improve the accuracy of audio detection.
[0004] In a first aspect, embodiments of the present disclosure provide an audio detection model training method, including:
[0005] Obtain audio sample data; wherein each of the audio sample data includes at least two types of audio sample sub-data;
[0006] Extract multi-modal audio features from the audio sample data through a multi-modal joint audio detection model to obtain multi-modal audio feature sample data;
[0007] Train the multi-modal joint audio detection model according to the multi-modal audio feature sample data.
[0008] In a second aspect, embodiments of the present disclosure provide an audio detection method, including:
[0009] Obtain audio data to be detected; wherein each of the audio data to be detected includes at least two types of audio sub-data to be detected;
[0010] Input the audio data to be detected into the multi-modal joint audio detection model for audio detection to obtain an audio detection result of the audio data to be detected;
[0011] Wherein, the multi-modal joint audio detection model is trained through the audio detection model training method described in the first aspect.
[0012] In a third aspect, embodiments of the present disclosure provide an audio detection model training device, including:
[0013] An audio sample data acquisition module, configured to obtain audio sample data; wherein each of the audio sample data includes at least two types of audio sample sub-data;
[0014] A multi-modal audio feature extraction module, configured to extract multi-modal audio features from the audio sample data through a multi-modal joint audio detection model to obtain multi-modal audio feature sample data;
[0015] A multi-modal joint audio detection model training module, configured to train a multi-modal joint audio detection model according to the multi-modal audio feature sample data.
[0016] Fourthly, an embodiment of the present disclosure provides an audio detection device, including:
[0017] A to-be-detected audio data acquisition module, configured to acquire to-be-detected audio data; wherein, each piece of the to-be-detected audio data includes at least two types of to-be-detected audio sub-data;
[0018] An audio detection result acquisition module, configured to input the to-be-detected audio data into a multi-modal joint audio detection model for audio detection to obtain an audio detection result of the to-be-detected audio data;
[0019] Wherein, the multi-modal joint audio detection model is trained by the audio detection model training device described in the third aspect.
[0020] Fifthly, an embodiment of the present disclosure provides an electronic device, including:
[0021] At least one processor; and
[0022] A memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the audio detection model training method provided in the embodiment of the first aspect or the audio detection method provided in the embodiment of the second aspect.
[0024] Sixthly, an embodiment of the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the audio detection model training method provided in the embodiment of the first aspect or the audio detection method provided in the embodiment of the second aspect.
[0025] Seventhly, an embodiment of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the audio detection model training method provided in the embodiment of the first aspect or the audio detection method provided in the embodiment of the second aspect is implemented.
[0026] In an embodiment of the present disclosure, a multi-modal joint audio detection model extracts multi-modal audio features from audio sample data of each acquired sample including at least two sub-data of audio samples, obtaining multi-modal audio feature sample data, and training the multi-modal joint audio detection model according to the extracted multi-modal audio feature sample data, so as to perform audio detection on the acquired audio data to be detected by using the multi-modal joint audio detection model, solving the problem of low accuracy of audio detection existing in the existing audio detection methods that only perform audio detection based on a single sample type or a single feature, thereby improving the accuracy of audio detection.
[0027] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0029] Figure 1 is a flowchart of a method for training an audio detection model provided by an embodiment of the present disclosure;
[0030] Figure 2 is a flowchart of a method for training an audio detection model provided by an embodiment of the present disclosure;
[0031] Figure 3 is a schematic diagram of the effect of noise interference data in audio data provided by an embodiment of the present disclosure;
[0032] Figure 4 is a schematic diagram of the effect of offset audio data in audio data provided by an embodiment of the present disclosure;
[0033] Figure 5 is a schematic diagram of the effect of a second segmented audio data provided by an embodiment of the present disclosure;
[0034] Figure 6 is a schematic diagram of the effect of a time-frequency diagram provided by an embodiment of the present disclosure;
[0035] Figure 7 is a schematic diagram of the training process of a multi-modal joint audio detection model provided by an embodiment of the present disclosure;
[0036] Figure 8 is a flowchart of an audio detection method provided by an embodiment of the present disclosure;
[0037] Figure 9 is a structural diagram of an audio detection model training device provided by an embodiment of the present disclosure;
[0038] Figure 10It is a structural diagram of an audio detection device provided by an embodiment of the present disclosure;
[0039] Figure 11 It is a schematic structural diagram of an electronic device for implementing the audio detection model training method or the audio detection method of the embodiments of the present disclosure. Detailed implementation manners
[0040] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0041] In one example, Figure 1 It is a flowchart of an audio detection model training method provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of training a multi-modal joint audio detection model using multi-modal audio features of audio sample data including multiple sub-samples. This method can be executed by an audio detection model training device, which can be implemented in software and / or hardware and is generally integrated in an electronic device. The electronic device can be a terminal device or a server device. The embodiments of the present disclosure do not limit the type of the electronic device. Correspondingly, as Figure 1 shown, the method includes the following operations:
[0042] S110. Obtain audio sample data, where each of the audio sample data includes at least two types of audio sample sub-data.
[0043] Among them, the audio sample data can be any type of data that needs to be subjected to audio detection. Exemplarily, the audio sample data can be the voice data of a user, or the audio data in a video, or the audio data collected by an audio device, such as the audio data collected when a certain device is working, etc. The embodiments of the present disclosure do not limit the audio type and audio content of the audio sample data. The audio sample sub-data can be the sub-sample data included in an audio sample data. Each audio sample sub-data can be the sample data corresponding to one state or type of the corresponding audio.
[0044] In the embodiments of the present disclosure, before training the multi-modal joint audio detection model, it is first necessary to obtain audio sample data. It can be understood that the multi-modal joint audio detection model can be applied to various different audio detection scenarios. Therefore, it is necessary to obtain audio sample data according to specific audio detection requirements. In order to further improve the diversity of the sample data, the audio sample data can include multiple audio sample sub-data.
[0045] Exemplarily, if it is necessary to detect the user's voice data, a large number of male voice sub-samples and female voice sub-samples can be collected as audio sample sub-data. That is, the male voice sample and the female voice sample are respectively used as a group of audio sample data. For example, the male voice sample and the female voice sample are both "Hello" to detect the user's voice "Hello". Alternatively, the voice word samples of the same user in a normal state and a sick state can be collected as the audio sample data corresponding to the user. If it is necessary to detect the working audio of the device to determine whether the device is in a normal working state and an abnormal working state, a large number of working audio data of the device in different working states can be collected as audio sample data. For example, the audio sample sub-data of a device in the initialization state, the running state and the shutdown state can be collected respectively to constitute the audio sample data of the working audio of the device, so as to comprehensively consider the comprehensive working state of the device to determine whether the device is abnormal.
[0046] S120. Extract multimodal audio features from the audio sample data using a multimodal joint audio detection model to obtain multimodal audio feature sample data.
[0047] Among them, the multimodal audio features may include a variety of different audio features. Different audio features may be audio features of different dimensions, or different types of audio features under the same dimension. Audio features of a single dimension may be, for example, frequency domain features, time domain features, acoustic features, and short-time frame features of audio, wherein the audio features of a single dimension may also be subdivided into a plurality of different types of audio features, for example, frequency domain features may include total energy, frequency center of gravity, frequency spectrum, root mean square frequency, kurtosis, skewness, kurtosis, and acoustic features may include MFCC (Mel-FrequencyCepstral Coefficients) and MEL filter group features, etc. Correspondingly, the multimodal audio features may be a combination of multiple different single-dimensional audio features, or may also be a combination of different types of audio features in a certain dimensional audio feature. For example, it may be a combination of time domain features and frequency domain features, or may also be a combination of multiple different forms of time domain features, frequency domain features, and acoustic features, or may also be a combination of different types of frequency domain features. Preferably, the multimodal audio features may include all types of audio features in all dimensions. As long as different audio features can be included to form a multimodal audio feature, the embodiment of the present disclosure does not limit the feature dimension and feature type of the multimodal audio feature. The multimodal audio feature sample data is sample data composed of multiple multimodal audio features.
[0048] Correspondingly, after obtaining the audio sample data, multi-modal audio features can be extracted from the obtained audio sample data through a multi-modal joint audio detection model. It can be understood that the richer the feature dimension and feature type of the multi-modal audio features, the richer the feature types included in the multi-modal audio feature sample data composed of the multi-modal audio features. Therefore, the detection accuracy and accuracy of the multi-modal joint audio detection model trained according to the multi-modal audio feature sample data are higher.
[0049] Optionally, extracting multi-modal audio features from the audio sample data may be to extract multi-modal audio features from each audio sample sub-data included in each audio sample data respectively, and then splice the multi-modal audio features extracted from the audio sample sub-data of each audio sample to obtain the multi-modal audio feature sample data corresponding to the audio sample.
[0050] S130. Train a multi-modal joint audio detection model according to the multi-modal audio feature sample data.
[0051] Among them, the multi-modal joint audio detection model can be a model capable of extracting and recognizing multi-modal audio feature sample data. Optionally, the multi-modal joint audio detection model can be composed of multiple different audio detection models, and each audio detection model can detect and recognize one type of audio feature sample data in the multi-modal audio feature sample data. Alternatively, the multi-modal joint audio detection model can also be composed of only one audio detection model, and this audio detection model can simultaneously extract and recognize multiple different types of audio feature sample data. As long as the extraction and recognition of the multi-modal audio feature sample data can be completed, the embodiments of the present disclosure do not limit the model composition and model architecture of the multi-modal joint audio detection model.
[0052] In the prior art, an audio detection model is trained by using a single sample type or audio feature. For example, when detecting the working state of a device, only the audio data of the device in normal and abnormal states during operation is often collected as sample data. However, abnormalities may occur in the device during the initialization configuration state and the shutdown state. Therefore, it is difficult to meet the detection requirements by using only the audio sample data in a single state for audio detection, and the detection accuracy is not high. Another example is to use only the MFCC acoustic feature of the audio for audio classification detection and recognition, or to use only the frequency range feature of the audio in the frequency domain for audio classification detection and recognition, etc. It can be understood that a single audio feature can only reflect the simple feature attributes of the audio. When different types of audio have the same feature attributes in this audio feature, the audio detection model cannot effectively detect and recognize different audio types. Exemplarily, when classifying and detecting male and female audio, if only the frequency of the audio is used for audio detection, neutral audio may be detected as both male and female audio, resulting in low audio detection accuracy. Or, when classifying and detecting normal and abnormal audio, if only the MFCC of the audio is used for audio detection, it is easy to detect normal audio as abnormal audio or abnormal audio as normal audio, resulting in low audio detection accuracy.
[0053] Therefore, expanding the sample type of the audio sample data first can enrich the diversity of the audio sample data, ensure the comprehensiveness of the model training information, and thus improve the accuracy of audio detection by the multi-modal joint audio detection model. At the same time, extracting the multi-modal audio feature sample data composed of the multi-modal audio features of the audio sample data can include the combination of complex feature attributes of the audio sample data, and the feature representation is more comprehensive and accurate. Thus, the useful information of the audio sample data can be comprehensively extracted, and then the multi-modal joint audio detection model can be trained according to the comprehensive useful information, which can improve the sensitivity of the multi-modal joint audio detection model to various audio features, maximize the avoidance of detection errors, and thus improve the accuracy of audio detection by the multi-modal joint audio detection model.
[0054] In the embodiment of the present disclosure, the multi-modal joint audio detection model extracts multi-modal audio features from the audio sample data in which each sample includes at least two audio sample sub-data to obtain multi-modal audio feature sample data, and trains the multi-modal joint audio detection model according to the extracted multi-modal audio feature sample data, so as to use the multi-modal joint audio detection model to perform audio detection on the obtained audio data to be detected, and solve the problem of low audio detection accuracy existing in the existing audio detection method that only performs audio detection based on a single sample type or a single feature, thereby improving the accuracy of audio detection.
[0055] In one example,Figure 2 Figure 2 is a flowchart of a method for training an audio detection model provided by an embodiment of the present disclosure. Based on the technical solutions of the above embodiments, the embodiment of the present disclosure has been optimized and improved, and various specific and optional implementation manners for obtaining audio sample data, extracting multi-modal audio features from the audio sample data, and training a multi-modal joint audio detection model according to the multi-modal audio feature sample data are given.
[0056] In an optional embodiment of the present disclosure, the audio sample data may include motor operation audio sample data, and the motor operation audio sample data may include motor forward rotation audio sample sub-data and motor reverse rotation audio sample sub-data.
[0057] Among them, the motor operation audio sample data is the audio sample data collected when the motor is operating. The motor can be any type of motor, such as an automotive seat motor or an elevator motor, etc. The embodiment of the present disclosure does not limit the type of the motor. It can be understood that the operation modes of the motor include two forms: forward rotation and reverse rotation. The motor forward rotation audio sample sub-data is the audio sample data collected when the motor rotates forward, and the motor reverse rotation audio sample sub-data is the audio sample data collected when the motor rotates in reverse. It can be understood that the motor forward rotation audio sample sub-data and the motor reverse rotation audio sample sub-data are the audio sample data collected by the same motor when it rotates forward and in reverse.
[0058] When evaluating the quality of a motor, the motor operation audio can be detected, and the motor quality can be judged according to whether there is an abnormality in the motor operation audio. In the prior art, often when the motor is operating, experienced inspectors judge whether the motor quality is qualified by manually listening to the sound. This way of manually detecting whether there is an abnormality in the motor operation audio has extremely high requirements for the inspectors. In a large factory, a production line can produce thousands of motor motors in a day. Using the method of manually detecting the motor operation audio requires a large amount of manpower for quality judgment, and the auditory sensitivities of different human ears are different, and the quality judgment is greatly affected by subjective factors. Therefore, there are often some motors with quality disputes, resulting in inconsistent motor quality evaluation criteria and easy fluctuations in the yield rate of motors.
[0059] Therefore, in order to improve the efficiency of audio detection for motor operation audio sample data, an audio detection model can be pre-trained, and the audio detection model can be used to replace the manual detection method to automatically judge the motor operation audio sample data, so as to automatically determine whether the motor operation audio sample data is normal data or abnormal data, and realize the automatic and high-efficiency detection of the motor operation audio sample data.
[0060] As Figure 2 shown, a method for training an audio detection model includes:
[0061] S210. Obtain audio sample data; each of the audio sample data includes at least two types of audio sample sub-data.
[0062] Correspondingly, S210 may specifically include the following operations:
[0063] S211. Obtain original audio data; each of the original audio data includes at least two types of original audio sub-data.
[0064] Among them, the original audio data may be audio data that has not been preprocessed. The original audio sub-data may be different types of sub-data that make up the original audio data.
[0065] Taking a motor as an example, the original audio data may be the audio data originally collected by an audio acquisition device when the motor is running. Since some motors have no abnormalities during forward rotation and have abnormalities during reverse rotation. Therefore, in order to achieve a comprehensive detection of the motor, the original audio data collected for one motor may include two original audio sub-data of the motor during forward rotation and reverse rotation.
[0066] S212. Perform data preprocessing on the original audio data to obtain preprocessed audio data.
[0067] In an optional embodiment of the present disclosure, the performing data preprocessing on the original audio data may include: performing normalization processing on the original audio data to obtain normalized audio data; performing truncation processing on the normalized audio data according to the noise interference data in the normalized audio data to obtain first segmented audio data; deleting the offset audio data in the first segmented audio data to obtain second segmented audio data; performing standardization processing on the data length of the second segmented audio data to obtain the preprocessed audio data.
[0068] Among them, the normalized audio data may be the audio data obtained by performing normalization processing on the original audio data. The first segmented audio data may be the segmented audio data obtained by deleting the noise interference data from the normalized audio data. The second segmented audio data may be the segmented audio data obtained by deleting the offset audio data from the first segmented audio data. The offset audio data is also the audio data that has shifted.
[0069] Optionally, if the original audio data is directly collected by an audio acquisition device through an AD (Analog-to-Digital Convert) chip, the original audio data is digital audio data. Therefore, normalization operation can be performed on the original audio data. According to the data acquisition characteristics, the formula: X = x / 32767 can be used to standardize the range of the original audio data to the interval of [-1, 1]. Where x represents the original audio data and X represents the normalized audio data.
[0070] Affected by the installation of the audio acquisition device and the structure of the device itself, when collecting audio data, some noise will inevitably interfere with the data, such as Figure 3 the noise interference data shown on both sides. Therefore, statistical methods can be used to dynamically calculate the noise interference data caused by acquisition in the normalized audio data, and a truncation operation is performed on the noise interference data, that is, the noise interference data in the normalized audio data is deleted. After deleting the noise interference data, the normalized audio data will be divided into the first segmented audio data in the form of multiple segments.
[0071] In addition, affected by the installation position of the sensor in the audio acquisition device, the audio data will be offset, such as Figure 4 shown. Therefore, if it is determined that there is offset audio data in the first segmented audio data, a high-order processing method can be used to delete the offset audio data to obtain the second segmented audio data as shown in Figure 5 . As shown in Figure 5 , the noise interference and offset abnormal audio data segments have been removed from the second segmented audio data.
[0072] Since the lengths of the original audio data are different, the data lengths of the second segmented audio data obtained after the original audio data undergoes noise stage and offset deletion processing may be different. When using an audio detection model for audio detection, the audio sample data often needs to have a unified data length. Therefore, in order to meet the input requirements of the audio detection model, the data lengths of the second segmented audio data can be standardized, and all the second segmented audio data can be regularized to a unified data length to obtain preprocessed audio data.
[0073] Through the above technical solutions, by normalizing the original audio data, the problem of numerical calculation errors in feature extraction of audio data can be prevented, the convergence speed of the multi-modal joint audio detection model can be accelerated, and the accuracy of the multi-modal joint audio detection model can be improved. Deleting the noise interference data and offset audio data of the audio data can avoid the interference and influence of abnormal audio generated by non-subjective factors on the audio data, thereby improving the accuracy and accuracy of the multi-modal joint audio detection model.
[0074] In an alternative embodiment of the present disclosure, the step of normalizing the data length of the second segmented audio data to obtain the preprocessed audio data includes: when it is determined that the data length of the second segmented audio data is less than the standard length of the segmented audio, determining the padding length of the second segmented audio data according to the data length of the second segmented audio data and the standard length of the segmented audio; collecting the collected audio data of the padding length from the second segmented audio data; and padding the second segmented audio data according to the collected audio data to obtain the preprocessed audio data; or, when it is determined that the data length of the second segmented audio data is greater than the standard length of the segmented audio, intercepting the intercepted audio data of the standard length of the segmented audio from the second segmented audio data; and using the intercepted audio data as the preprocessed audio data.
[0075] Among them, the standard length of the segmented audio can be a unified standard length preset for the segmented audio. The padding length can be the data length required to pad the second segmented audio data to the standard length of the segmented audio. The collected audio data can be the audio data collected from the second segmented audio data, and its data length is the padding length. The intercepted audio data is also the audio data intercepted from the second segmented audio data, and its data length is the standard length of the segmented audio.
[0076] When normalizing the data length of the second segmented audio data, the data length of the second segmented audio data can be analyzed first. If it is determined that the data length of the second segmented audio data is less than the standard length of the segmented audio, the padding length of the second segmented audio data can be calculated according to the difference between the standard length of the segmented audio and the data length of the second segmented audio data, and then the collected audio data of the padding length can be collected from the second segmented audio data, so as to pad the second segmented audio data according to the collected audio data to obtain the preprocessed audio data. Optionally, the collected audio data can be inserted at any position of the second segmented audio data to achieve the padding process. Among them, the method of collecting the collected audio data of the padding length can be random collection or collection according to a preset data collection rule, and the embodiments of the present disclosure do not limit this. Among them, padding the second segmented audio data by randomly collecting audio data can ensure the feature consistency of the second segmented audio data, avoid adding other audio data to cause adverse interference to the second segmented audio data, and thus ensure the accuracy and reliability of the audio sample data.
[0077] Correspondingly, if it is determined that the data length of the second segmented audio data is greater than the segmented audio standard length, the intercepted audio data of the segmented audio standard length can be intercepted from the second segmented audio data. Exemplarily, the data at both ends of the second segmented audio data can be deleted, and the audio data of the segmented audio standard length at the middle position can be retained as the intercepted audio data to maximize the retention of useful information in the audio data, and the intercepted audio data is used as the preprocessed audio data. If the useful information of the second segmented audio data is evenly distributed, the second segmented audio data can be directly divided into at least two segment data, and each segment data is respectively padded to obtain a plurality of preprocessed audio data.
[0078] Taking the motor operation audio sample data as an example, a set of original forward rotation audio sub-data and reverse rotation audio sub-data corresponding to a motor can be simultaneously subjected to the above preprocessing process. That is, after the original audio data of a motor is preprocessed, two preprocessed audio data, namely the preprocessed audio data of the motor forward rotation and the preprocessed audio data of the motor reverse rotation, can be obtained correspondingly.
[0079] S213. Obtain the audio sample label of the preprocessed audio data.
[0080] Among them, the audio sample label is also the label for marking the preprocessed audio data. Exemplarily, the audio sample label can be one or more groups of labels, such as labels for male and female, normal and abnormal, etc. The embodiments of the present disclosure do not limit the label type and the number of labels of the audio sample label.
[0081] In an optional embodiment of the present disclosure, the obtaining the audio sample label of the preprocessed audio data may include: obtaining the label judgment result of the preprocessed audio data; determining the audio sample label matched by the preprocessed audio data according to the label judgment result of the preprocessed audio data.
[0082] S214. Mark the preprocessed audio data according to the audio sample label to obtain the audio sample data.
[0083] Taking the audio sample data of motor operation as an example, after the original audio data of motor operation is preprocessed, the preprocessed audio data can be judged by manual detection to generate the label judgment result of normal or abnormal for the preprocessed audio data, and the preprocessed audio data can be marked according to the label judgment result. Optionally, one label can be correspondingly marked for one preprocessed audio data of motor operation. Exemplarily, when all the forward rotation audio sample sub-data and reverse rotation audio sample sub-data in the preprocessed audio sample data of motor operation are in the normal state, the preprocessed audio sample data of motor operation can be marked as normal. When at least one sub-sample data of the forward rotation audio sample sub-data and reverse rotation audio sample sub-data in the preprocessed audio sample data of motor operation is in the abnormal state, the preprocessed audio sample data of motor operation can be marked as abnormal. Optionally, the form of the audio sample label can be converted into the one-hot form for subsequent model training. It can be understood that the audio sample data can include one or more audio sample labels.
[0084] In the above technical solution, by performing data preprocessing on the original audio data and using labels to mark and process the preprocessed audio data, the complete acquisition and processing of audio sample data are realized, and the availability and accuracy of the audio sample data are ensured.
[0085] S220. Extract multi-modal audio features from the audio sample data through a multi-modal joint audio detection model to obtain multi-modal audio feature sample data.
[0086] In an optional embodiment of the present disclosure, the audio sample data includes first audio sample sub-data and second audio sample sub-data.
[0087] Among them, the first audio sample sub-data and the second audio sample sub-data may be audio sample data of different states of the same detection object. Taking the motor as the detection object as an example, the first audio sample sub-data may be the forward rotation audio sample sub-data of the motor, and the second audio sample sub-data may be the reverse rotation audio sample sub-data of the motor.
[0088] Correspondingly, S220 may specifically include the following operations:
[0089] S221. Extract a set number of single-modal audio features from the first audio sample sub-data through the multi-modal joint audio detection model.
[0090] Among them, the set number can be set according to actual needs, such as 3, 5 or 6, etc., and the embodiments of the present disclosure do not limit this. The single-modal audio feature may be an audio feature of one dimension.
[0091] S222. Extract the set number of unimodal audio features from the second audio sample sub-data through the multi-modal joint audio detection model.
[0092] Among them, the unimodal audio features can include but are not limited to spectrogram features, MFCC features, Mel MEL filter bank features, time-domain mathematical statistical features, frequency-domain mathematical statistical features, and short-time frame features.
[0093] In the embodiments of the present disclosure, it is necessary to extract the set number of unimodal audio features from the first audio sample sub-data and the second audio sample sub-data at the same time. Each unimodal audio feature can also include multiple different types of audio features.
[0094] Exemplarily, when extracting spectrogram features, short-time Fourier transform (Short-time Fourier Transform, or Short-term Fourier Transform, STFT) can be performed on the first audio sample sub-data and the second audio sample sub-data to obtain the spectrogram corresponding to the signal as Figure 6 shown. The Hamming window is selected for the Fourier transform, the window length can be 512, and the step size can be 256. The spectrogram represents the energy representation of the signal in the time domain and the frequency domain, and is in the form of a two-dimensional matrix. This feature contains the most information in the first audio sample sub-data and the second audio sample sub-data.
[0095] Optionally, if audio detection needs to refer to manual detection factors, the MFCC feature can be used as one of the unimodal audio features. The extraction process of the MFCC feature can be specifically: continuous speech - pre-emphasis - windowing and framing - Fourier transform - Mel filter bank - logarithmic operation - dynamic differential parameter extraction - feature vector. The dimension of the MFCC feature can be [1×n], where n is the number of filters in the Mel filter bank. Optionally, the value of n can be 40.
[0096] Since the MFCC feature only considers dynamic differential parameters and there is a certain loss of original information, therefore, optionally, the Mel MEL filter bank feature can also be used as one of the unimodal audio features. The Mel MEL filter bank feature can also be in the form of a two-dimensional matrix, and the dimension of the feature can be [t×n], where t is the same as the number of columns of the spectrogram, and n can be the number of filters in the Mel filter bank.
[0097] Optionally, the time-domain mathematical statistical feature may be a mathematical feature of the first audio sample sub-data and the second audio sample sub-data on a time scale. Optionally, the time-domain mathematical statistical feature may include, but is not limited to, mean, peak, rectified mean, root mean square, kurtosis, skewness, kurtosis, waveform factor, impulse factor, root amplitude, margin factor, and peak factor. The dimension of the time-domain mathematical statistical feature may be [1×m], where m represents different types of mathematical features counted on a time scale.
[0098] Optionally, the frequency domain mathematical statistical features may be mathematical characteristics of the audio signal mapped to the frequency domain after fast Fourier transforming the first audio sample sub-data and the second audio sample sub-data, and then statistically analyzing the transformed audio signal on the frequency scale. Optionally, the frequency domain mathematical statistical features may include but are not limited to total energy, frequency center of gravity, frequency spectrum, root mean square frequency, kurtosis, skewness, kurtosis, waveform factor, pulse factor, root amplitude, margin factor, and peak factor. The dimension of the time domain mathematical statistical features may be [1×M], where M represents different types of mathematical features statistically analyzed on the time scale. The dimensions of the time domain mathematical statistical features and the frequency domain mathematical statistical features may be the same or different, and the embodiments of the present disclosure do not limit this.
[0099] Optionally, extracting short-time frame features from the audio sample data may include: performing short-time slicing processing on the first audio sample sub-data and the second audio sample sub-data, thereby dividing the first audio sample sub-data and the second audio sample sub-data into frame segments of the same size, and then performing separate calculations on each frame segment, and finally integrating the results of all frame segments, for example, calculating the average value of feature quantities such as energy, amplitude, zero-crossing rate, autocorrelation, and amplitude difference of all frame segments to obtain short-time energy, short-time average amplitude, short-time average zero-crossing rate, short-time autocorrelation, short-time average amplitude difference, etc. Optionally, the dimension of the short-time frame feature may be [1×5].
[0100] S223: Splice the unimodal audio features of the first audio sample sub-data and the unimodal audio features of the second audio sample sub-data to obtain spliced unimodal audio feature sample data.
[0101] S224. Construct the multimodal audio feature sample data according to the concatenated single-modal audio feature sample data.
[0102] The spliced unimodal audio feature sample data may be feature data obtained by splicing and fusing unimodal audio features of the same dimension of the first audio sample sub-data and the second audio sample sub-data.
[0103] Since the first audio sample sub-data and the second audio sample sub-data belong to the same-source audio sample data, after extracting a set number of unimodal audio features from the first audio sample sub-data and the second audio sample sub-data respectively, the unimodal audio features of the first audio sample sub-data and the unimodal audio features of the second audio sample sub-data can be concatenated and fused to obtain the concatenated unimodal audio features corresponding to each unimodal audio feature. Each concatenated unimodal audio feature can constitute the multimodal audio feature sample data of the audio sample data.
[0104] Optionally, the concatenation and fusion of the unimodal audio features of the first audio sample sub-data and the unimodal audio features of the second audio sample sub-data can be to concatenate each unimodal audio feature of the first audio sample sub-data and the second audio sample sub-data to obtain the concatenated unimodal audio sample features. For example, the time-frequency map features of the first audio sample sub-data and the second audio sample sub-data can be concatenated, and the time-domain mathematical statistical features of the first audio sample sub-data and the second audio sample sub-data can be concatenated, and so on until the concatenation and fusion processing of all multimodal audio features are completed.
[0105] Through the above technical solution, by concatenating the unimodal audio features of the audio sample data including the first audio sample sub-data and the second audio sample sub-data, the audio feature extraction processing of the audio data generated by the detection object in multiple different states can be realized, enriching the data content and feature representation of the audio sample data.
[0106] S230. Train a multimodal joint audio detection model according to the multimodal audio feature sample data.
[0107] In an optional embodiment of the present disclosure, the multimodal joint audio detection model includes a result summary model and a set number of audio detection models.
[0108] Among them, the audio detection model can identify and detect one type of unimodal audio feature, and the number of audio detection models is the same as the number of unimodal audio features. The result summary model can summarize and statistically analyze the audio detection results of each audio detection model to obtain the final audio detection result. Optionally, the architectures of each audio detection model can be different, and the architectures of the audio detection model and the result summary model can also be different.
[0109] Correspondingly, S230 can specifically include the following operations:
[0110] S231. Input each concatenated unimodal audio feature sample data of the multimodal audio feature sample data into each audio detection model of the multimodal joint audio detection model to obtain the unimodal audio detection results of each audio detection model.
[0111] Among them, the single-modal audio detection result can be the result of the audio detection model detecting each spliced single-modal audio feature sample data.
[0112] Figure 7 It is a schematic diagram of the training process of a multi-modal joint audio detection model provided by an embodiment of the present disclosure. In a specific example, as Figure 7 shown, taking the motor forward rotation audio sample sub-data (i.e., forward rotation signal) and the motor reverse rotation audio sample sub-data (i.e., reverse rotation signal) of the motor as the audio sample data as an example, assuming that a total of 6 different types of single-modal audio features can be extracted from the audio sample data, the multi-modal joint audio detection model can include a total of 7 models, among which 6 models are audio detection models, that is, as Figure 7 shown, 6 neural networks, and one is a result summary model, that is, as Figure 7 shown, the result voting neural network. It can be understood that for different feature quantities, the corresponding network modules are also different. For picture-like features such as time-frequency map features and Mel filter bank features, a combination of modules such as two-dimensional convolution, Batch Normalization, max pooling, ReLU (Rectified Linear Units, activation function), and fully connected can be used to construct the corresponding audio detection model. For one-dimensional data-like features such as MFCC features, time-domain mathematical statistical features, frequency-domain mathematical statistical features, and short-time frame features, a combination of modules such as one-dimensional convolution and Batch Normalization, max pooling, ReLU, and fully connected can be used to construct the corresponding audio detection model.
[0113] Since each spliced single-modal audio feature sample data (the feature quantities of forward rotation and reverse rotation are spliced together) can include feature quantities of different dimensions, therefore, each spliced single-modal audio feature sample data can be respectively input into the corresponding audio detection model. Each audio detection model can calculate the individual feature quantity and obtain a separate 1*2-dimensional audio detection result as the single-modal audio detection result. The first component result in the single-modal audio detection result can represent the probability of normal audio, and the second component result can represent the probability of abnormal audio.
[0114] S232. Splice the single-modal audio detection results of each of the audio detection models to obtain a spliced single-modal audio detection result.
[0115] Correspondingly, after each audio detection model outputs the single-modal audio detection result, the single-modal audio detection results can be spliced to obtain a spliced single-modal audio detection result. It can be understood that the dimension of the spliced single-modal audio detection result is the same as that of the single-modal audio detection result.
[0116] S233. Input the spliced unimodal audio detection result into the result summarization model of the multimodal joint audio detection model to obtain the target audio detection result.
[0117] Among them, the target audio detection result is also the final detection result obtained by the result summarization model through summarizing and statistically analyzing each spliced unimodal audio detection result.
[0118] S234. Calculate the model loss according to the target audio detection result, and update the model parameters of the multimodal joint audio detection model according to the model loss.
[0119] Among them, the model loss can be the loss of each model in the multimodal joint audio detection model.
[0120] Exemplarily, as Figure 7 shown, the spliced unimodal audio detection results of all audio detection models can be input into the final result summarization model, and the result summarization model performs integrated calculation on the spliced unimodal audio detection results to obtain the final target audio detection result. Further, compare the target audio detection result with the label in the audio sample data, and calculate the loss of each model by using the comparison result and the loss function, so as to update the model parameters of each model in the multimodal joint audio detection model simultaneously according to the model loss.
[0121] Optionally, the loss function can adopt the cross-entropy loss, and its expression can be: Among them, L cross (Y, P) represents the model loss, Y represents the label data of the audio sample data, P represents the target audio detection result, n represents the number of training samples, k represents the number of training sample categories, y i,k represents the true label of the data set, and p i,k represents the model output value. At the same time, the stochastic gradient descent method can be used to minimize the loss function to realize the parameter update of all models in the multimodal joint audio detection model.
[0122] In the above technical solution, by using each audio detection model of the multimodal joint audio detection model to respectively identify and calculate each spliced unimodal audio feature sample data, the accuracy of audio detection of the spliced unimodal audio feature sample data can be improved, and then integrated calculation is performed on the spliced unimodal audio detection results obtained by splicing the unimodal audio detection results of each audio detection model, so as to accurately determine the final audio detection result by comprehensively combining the audio detection results of various different types of audio features.
[0123] In an alternative embodiment of the present disclosure, the obtaining of the target audio detection result may include: determining, by the result aggregation model, the number of the first audio detection results and the number of the second audio detection results according to the spliced unimodal audio detection results; and determining, by the result aggregation model, the target audio detection result according to the number of the first audio detection results, the number of the second audio detection results, and the weights of the spliced unimodal audio detection results.
[0124] Wherein, the first audio detection result and the second audio detection result may be two different types of detection results, such as two results of normal and abnormal.
[0125] It can be understood that the proportions of useful information of audio that different audio features can contain are also different. Therefore, corresponding weights can be set for different spliced unimodal audio detection results respectively to represent the importance of each spliced unimodal audio detection result, so as to clarify the influence degree of each audio feature on the audio detection result and improve the accuracy of audio detection. Optionally, the weight of the spliced unimodal audio detection result corresponding to the time-frequency graph feature is the highest, the weight value of the spliced unimodal audio detection result corresponding to the Mel filter bank feature ranks second, the weight value of the spliced unimodal audio detection result corresponding to the MFCC feature ranks third, and the weights of the spliced unimodal audio detection results corresponding to other types of audio features can be the smallest.
[0126] Therefore, the result aggregation model can consider two factors, namely the number of the first audio detection results, the number of the second audio detection results, and the weights of the spliced unimodal audio detection results, to determine the target audio detection result. Exemplarily, if the spliced unimodal audio detection results corresponding to the time-frequency graph feature and the time-domain mathematical statistical feature are normal, and the rest of the spliced unimodal audio detection results are abnormal, the weights of the spliced unimodal audio detection results corresponding to the time-frequency graph feature and the time-domain mathematical statistical feature are 0.5 and 0.08 respectively, and the weights of the rest of the spliced unimodal audio detection results are 0.2, 0.1, 0.05, and 0.07 respectively, then the probability of the normal detection result is: 0.5 + 0.08 = 0.58, and the probability of the abnormal detection result is: 0.2 + 0.1 + 0.05 + 0.07 = 0.42. Then the probability of the target audio detection result being normal is 0.58, and the probability of being abnormal is 0.42.
[0127] The above technical solution can improve the accuracy of audio detection of the multi-modal joint audio detection model by forming a multi-modal joint audio detection model with multiple different models, extracting multi-modal audio features through the multi-modal joint audio detection model to obtain multi-modal audio feature sample data, and then training the multi-modal joint audio detection model according to the multi-modal audio feature sample data.
[0128] In an example,Figure 8 This is a flowchart of an audio detection method provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of using a multi-modal joint audio detection model for audio detection. This method can be executed by an audio detection device, which can be implemented in the form of software and / or hardware and is generally integrated in an electronic device. The electronic device can be a terminal device or a server device. The embodiment of the present disclosure does not limit the type of the electronic device. Correspondingly, as Figure 8 shown, this method includes the following operations:
[0129] S810. Obtain the audio data to be detected; wherein, each of the audio data to be detected includes at least two sub-audio data to be detected.
[0130] Among them, the audio data to be detected can be the data that needs to be subjected to audio detection. The sub-audio data to be detected can be the audio sub-data included in a piece of audio data to be detected. Each piece of audio sub-data can be the data corresponding to one state or type of the corresponding audio.
[0131] In an optional embodiment of the present disclosure, the audio sample data can include the audio data of the motor to be detected during operation, and the audio sample data of the motor to be detected during operation can include the sub-audio data of the motor to be detected during forward rotation and the sub-audio data of the motor to be detected during reverse rotation.
[0132] Among them, the audio data of the motor to be detected during operation is the audio data that needs to be detected collected when the motor is operating. The motor can be any type of motor, such as an automobile seat motor or an elevator motor, etc. The embodiment of the present disclosure does not limit the type of the motor. It can be understood that the operation modes of the motor include two forms: forward rotation and reverse rotation. The sub-audio data of the motor to be detected during forward rotation is the audio data that needs to be detected collected when the motor rotates forward, and the sub-audio data of the motor to be detected during reverse rotation is the audio data that needs to be detected collected when the motor rotates in reverse. It can be understood that the sub-audio data of the motor to be detected during forward rotation and the sub-audio data of the motor to be detected during reverse rotation are the audio data that needs to be detected collected when the same motor rotates forward and in reverse.
[0133] S820. Input the audio data to be detected into the multi-modal joint audio detection model for audio detection to obtain the audio detection result of the audio data to be detected.
[0134] Among them, the multi-modal joint audio detection model is trained by any of the above-mentioned audio detection model training methods.
[0135] Correspondingly, after obtaining the audio data to be detected, the multi-modal joint audio detection model can be used to extract multi-modal audio features from the audio data to be detected, and obtain the multi-modal audio feature data of the audio data to be detected. Further, after the multi-modal joint audio detection model extracts the multi-modal audio feature data, audio detection can be performed according to the multi-modal audio feature data to obtain the audio detection result of the audio data to be detected.
[0136] Taking the motor as the detection object as an example, the input of the audio data to be detected received by the multi-modal joint audio detection model is 2 motor operation audio data of forward rotation and reverse rotation. The multi-modal joint audio detection model can extract, identify and detect the multi-modal audio features of the input audio data to be detected, so as to directly give the probabilities of normal and abnormal of the motor corresponding to the audio data to be detected.
[0137] In the embodiments of the present disclosure, the multi-modal joint audio detection model extracts multi-modal audio features from the audio sample data of each sample including at least two audio sample sub-data, and obtains multi-modal audio feature sample data, so as to train the multi-modal joint audio detection model according to the extracted multi-modal audio feature sample data, thereby using the multi-modal joint audio detection model to perform audio detection on the audio data to be detected, and solving the problem of low audio detection accuracy existing in the existing audio detection methods that only perform audio detection based on a single sample type or a single feature, thereby improving the accuracy of audio detection.
[0138] In the technical solution of the present disclosure, the processing of collecting, storing, using, processing, transmitting, providing and disclosing the user's personal information (such as the user's voice information, etc.) all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0139] It should be noted that any permutation and combination of the technical features in the above embodiments also belong to the protection scope of the present disclosure.
[0140] In one example, Figure 9 is the structural diagram of an audio detection model training device provided by the embodiments of the present disclosure. The embodiments of the present disclosure are applicable to the situation of training a multi-modal joint audio detection model using the multi-modal audio features of audio sample data including multiple sub-samples. The device is implemented by software and / or hardware and is specifically configured in an electronic device. The electronic device can be a terminal device or a server device. The embodiments of the present disclosure do not limit the device type of the electronic device.
[0141] Such as Figure 9 An audio detection model training device 900 shown includes: an audio sample data acquisition module 910, a multi-modal audio feature extraction module 920, and a multi-modal joint audio detection model training module 930. Among them,
[0142] An audio sample data acquisition module 910 is configured to acquire audio sample data; wherein, each of the audio sample data includes at least two types of audio sample sub-data;
[0143] A multi-modal audio feature extraction module 920 is configured to extract multi-modal audio features from the audio sample data through a multi-modal joint audio detection model to obtain multi-modal audio feature sample data;
[0144] A multi-modal joint audio detection model training module 930 is configured to train the multi-modal joint audio detection model according to the multi-modal audio feature sample data.
[0145] In an embodiment of the present disclosure, the multi-modal joint audio detection model is used to extract multi-modal audio features from the acquired audio sample data in which each sample includes at least two types of audio sample sub-data to obtain multi-modal audio feature sample data, so as to train the multi-modal joint audio detection model according to the extracted multi-modal audio feature sample data, thereby using the multi-modal joint audio detection model to perform audio detection on the acquired audio data to be detected, and solving the problem of low audio detection accuracy existing in the existing audio detection method that only performs audio detection based on a single sample type or a single feature, thereby improving the accuracy of audio detection.
[0146] Optionally, the audio sample data acquisition module 910 is further configured to: acquire original audio data; wherein, each of the original audio data includes at least two types of original audio sub-data; perform data preprocessing on the original audio data to obtain preprocessed audio data; acquire audio sample labels of the preprocessed audio data; and label the preprocessed audio data according to the audio sample labels to obtain the audio sample data.
[0147] Optionally, the audio sample data acquisition module 910 is further configured to: perform normalization processing on the original audio data to obtain normalized audio data; perform truncation processing on the normalized audio data according to noise interference data in the normalized audio data to obtain first segmented audio data; delete offset audio data in the first segmented audio data to obtain second segmented audio data; and perform standardization processing on the data length of the second segmented audio data to obtain the preprocessed audio data.
[0148] Optionally, the audio sample data acquisition module 910 is further configured to: when determining that the data length of the second segmented audio data is less than the segmented audio standard length, determine the padding length of the second segmented audio data according to the data length of the second segmented audio data and the segmented audio standard length; collect the collected audio data of the padding length from the second segmented audio data; and pad the second segmented audio data according to the collected audio data to obtain the preprocessed audio data; or, when determining that the data length of the second segmented audio data is greater than the segmented audio standard length, intercept the segmented audio standard length of intercepted audio data from the second segmented audio data; and use the intercepted audio data as the preprocessed audio data.
[0149] Optionally, the audio sample data includes first audio sample sub-data and second audio sample sub-data; the multi-modal audio feature extraction module 920 is further configured to: extract a set number of single-modal audio features from the first audio sample sub-data through the multi-modal joint audio detection model; extract the set number of single-modal audio features from the second audio sample sub-data through the multi-modal joint audio detection model; splice the single-modal audio features of the first audio sample sub-data and the single-modal audio features of the second audio sample sub-data to obtain spliced single-modal audio feature sample data; and constitute the multi-modal audio feature sample data according to each of the spliced single-modal audio feature sample data.
[0150] Optionally, the single-modal audio feature includes at least one of the following: time-frequency diagram feature, Mel cepstral coefficient MFCC feature, Mel MEL filter bank feature, time-domain mathematical statistical feature, frequency-domain mathematical statistical feature, and short-time segmented frame feature.
[0151] Optionally, the time-domain mathematical statistical feature includes at least one of the following: mean, peak value, rectified average value, root mean square, kurtosis, skewness, kurtosis, waveform factor, pulse factor, root mean square amplitude, margin factor, and peak factor; and / or, the frequency-domain mathematical statistical feature includes at least one of the following: total energy, frequency centroid, frequency spectrum line, root mean square frequency, kurtosis, skewness, kurtosis, waveform factor, pulse factor, root mean square amplitude, margin factor, peak factor; and / or, the short-time segmented frame feature includes short-time energy, short-time average amplitude, short-time average zero-crossing rate, short-time autocorrelation, and short-time average amplitude difference.
[0152] Optionally, the multi-modal joint audio detection model includes a result summarization model and a set number of audio detection models; the multi-modal joint audio detection model training module 930 is further configured to: input each spliced single-modal audio feature sample data of the multi-modal audio feature sample data into each audio detection model of the multi-modal joint audio detection model to obtain single-modal audio detection results of each audio detection model; splice the single-modal audio detection results of each audio detection model to obtain a spliced single-modal audio detection result; input the spliced single-modal audio detection result into the result summarization model of the multi-modal joint audio detection model to obtain a target audio detection result; calculate a model loss according to the target audio detection result, and update model parameters of the multi-modal joint audio detection model according to the model loss.
[0153] Optionally, the multi-modal joint audio detection model training module 930 is further configured to: determine the number of first audio detection results and the number of second audio detection results according to the spliced single-modal audio detection result through the result summarization model; determine the target audio detection result according to the number of first audio detection results and the number of second audio detection results, and the weights of the spliced single-modal audio detection results through the result summarization model.
[0154] Optionally, the audio sample data includes motor operation audio sample data, and the motor operation audio sample data includes motor forward rotation audio sample sub-data and motor reverse rotation audio sample sub-data.
[0155] The above audio detection model training device can execute the audio detection model training method provided by any embodiment of the present disclosure, and has corresponding function modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the audio detection model training method provided by any embodiment of the present disclosure.
[0156] In one example, Figure 10 is a structural diagram of an audio detection device provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of audio detection using a multi-modal joint audio detection model. The device is implemented by software and / or hardware and is specifically configured in an electronic device. The electronic device can be a terminal device or a server device. The embodiment of the present disclosure does not limit the device type of the electronic device.
[0157] As Figure 10 shown, an audio detection model training device 1000 includes: a to-be-detected audio data acquisition module 1010 and an audio detection result acquisition module 1020. Among them,
[0158] An audio data to be detected acquisition module 1010 is configured to acquire audio data to be detected; wherein, each of the audio data to be detected includes at least two types of audio sub-data to be detected;
[0159] An audio detection result acquisition module 1020 is configured to input the audio data to be detected into a multi-modal joint audio detection model for audio detection to obtain an audio detection result of the audio data to be detected;
[0160] Wherein, the multi-modal joint audio detection model is trained by the audio detection model training device described in any one of the above.
[0161] The above audio detection model training device can execute the audio detection model training method provided in any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution of the method. Technical details not described in detail in this embodiment can be referred to the audio detection model training method provided in any embodiment of the present disclosure.
[0162] In one example, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0163] Figure 11 A schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0164] As Figure 11 shown, the device 1100 includes a computing unit 1101, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0165] Multiple components in device 1100 are connected to I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0166] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the audio detection model training method or the audio detection method. For example, in some embodiments, the audio detection model training method or the audio detection method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the audio detection model training method or the audio detection method described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the audio detection model training method or the audio detection method in any other suitable manner (e.g., by means of firmware).
[0167] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, where the programmable processor can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0168] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0169] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0170] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0171] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such backend, middleware, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0172] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system, or a server combined with a blockchain.
[0173] In an embodiment of the present disclosure, a multi-modal joint audio detection model extracts multi-modal audio features from audio sample data of each sample obtained, where the audio sample data includes at least two sub-data of audio samples, to obtain multi-modal audio feature sample data, and trains the multi-modal joint audio detection model according to the extracted multi-modal audio feature sample data, so as to use the multi-modal joint audio detection model to perform audio detection on the to-be-detected audio data obtained, solving the problem of low audio detection accuracy existing in the existing audio detection methods that only perform audio detection based on a single sample type or a single feature, thereby improving the accuracy of audio detection.
[0174] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0175] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. An audio detection model training method, comprising: Obtaining audio sample data; wherein each of the audio sample data includes at least two types of audio sample sub-data; Extracting multi-modal audio features from the audio sample data through a multi-modal joint audio detection model to obtain multi-modal audio feature sample data; wherein the multi-modal audio features include multiple types of audio features in multiple dimensions; Training the multi-modal joint audio detection model according to the multi-modal audio feature sample data; Wherein the audio sample data includes first audio sample sub-data and second audio sample sub-data; the first audio sample sub-data and the second audio sample sub-data are audio sample data of different states of the same detection object; the detection object includes a motor; the first audio sample sub-data includes motor forward rotation audio sample sub-data, and the second audio sample sub-data includes motor reverse rotation audio sample sub-data; the extracting multi-modal audio features from the audio sample data through a multi-modal joint audio detection model to obtain multi-modal audio feature sample data includes: Extracting a set number of single-modal audio features from the first audio sample sub-data through the multi-modal joint audio detection model; Extracting the set number of single-modal audio features from the second audio sample sub-data through the multi-modal joint audio detection model; Concatenating the single-modal audio features of the first audio sample sub-data and the single-modal audio features of the second audio sample sub-data to obtain concatenated single-modal audio feature sample data; wherein the concatenated single-modal audio feature sample data is feature data obtained by concatenating and fusing the single-modal audio features of a certain same dimension of the first audio sample sub-data and the second audio sample sub-data; Constructing the multi-modal audio feature sample data according to each of the concatenated single-modal audio feature sample data.
2. The method according to claim 1, wherein, The obtaining audio sample data includes: Obtaining original audio data; wherein each of the original audio data includes at least two types of original audio sub-data; Performing data preprocessing on the original audio data to obtain preprocessed audio data; Obtaining audio sample labels of the preprocessed audio data; Marking the preprocessed audio data according to the audio sample labels to obtain the audio sample data.
3. According to the method of claim 2, the performing data preprocessing on the original audio data includes: Performing normalization processing on the original audio data to obtain normalized audio data; Performing truncation processing on the normalized audio data according to the noise interference data in the normalized audio data to obtain first segmented audio data; Deleting the offset audio data in the first segmented audio data to obtain second segmented audio data; Performing standardization processing on the data length of the second segmented audio data to obtain the preprocessed audio data.
4. The method according to claim 3, wherein The performing standardization processing on the data length of the second segmented audio data to obtain the preprocessed audio data includes: In the case where it is determined that the data length of the second segmented audio data is less than the segmented audio standard length, determining the padding length of the second segmented audio data according to the data length of the second segmented audio data and the segmented audio standard length; collecting collected audio data of the padding length from the second segmented audio data; and padding the second segmented audio data according to the collected audio data to obtain the preprocessed audio data; or When it is determined that the data length of the second segmented audio data is greater than the segmented audio standard length, intercepted audio data of the segmented audio standard length is intercepted from the second segmented audio data; and the intercepted audio data is used as the preprocessed audio data.
5. The method according to claim 1, wherein, The unimodal audio feature includes at least one of the following: a time-frequency graph feature, a Mel-frequency cepstral coefficient (MFCC) feature, a Mel-frequency MEL filter bank feature, a time-domain mathematical statistical feature, a frequency-domain mathematical statistical feature, and a short-time frame feature.
6. The method according to claim 5, wherein, The time domain mathematical statistical characteristics include at least one of the following: mean, peak, rectified mean, root mean square, kurtosis, skewness, kurtosis, waveform factor, pulse factor, root amplitude, margin factor and peak factor; and / or The frequency domain mathematical statistical features include at least one of the following: total energy, frequency center of gravity, frequency spectrum, root mean square frequency, kurtosis, skewness, kurtosis, shape factor, pulse factor, root amplitude, margin factor and peak factor; and / or The short-time framing feature includes at least one of the following: short-time energy, short-time average amplitude, short-time average zero-crossing rate, short-time autocorrelation and short-time average amplitude difference.
7. The method according to any one of claims 1 and 5-6, wherein The multimodal joint audio detection model includes a result aggregation model and a set number of audio detection models; The training of the multimodal joint audio detection model according to the multimodal audio feature sample data includes: Inputting each concatenated single-modal audio feature sample data of the multimodal audio feature sample data into each audio detection model of the multimodal joint audio detection model, respectively, to obtain a single-modal audio detection result of each audio detection model; Splicing the single-modal audio detection results of each of the audio detection models to obtain a spliced single-modal audio detection result; Inputting the spliced single-modal audio detection result into the result summary model of the multimodal joint audio detection model to obtain the target audio detection result; The model loss is calculated according to the target audio detection result, and the model parameters of the multimodal joint audio detection model are updated according to the model loss.
8. The method according to claim 7, wherein The obtaining of the target audio detection result includes: Determine the number of first audio detection results and the number of second audio detection results according to the spliced single-modal audio detection result by using the result summary model; The target audio detection result is determined by the result summary model according to the number of the first audio detection results and the number of the second audio detection results, and the weight of each of the spliced single-modal audio detection results.
9. An audio detection method, comprising: Acquire audio data to be detected; wherein the audio data to be detected includes at least two types of audio sub-data to be detected; Input the audio data to be detected into a multi-modal joint audio detection model for audio detection to obtain the audio detection result of the audio data to be detected; Among them, the multi-modal joint audio detection model is trained by the audio detection model training method described in any one of claims 1-8.
10. An audio detection model training device, comprising: An audio sample data acquisition module, configured to acquire audio sample data; among them, each of the audio sample data includes at least two types of audio sample sub-data; A multi-modal audio feature extraction module, configured to extract multi-modal audio features from the audio sample data through a multi-modal joint audio detection model to obtain multi-modal audio feature sample data; among them, the multi-modal audio features include multiple types of audio features in multiple dimensions; A multi-modal joint audio detection model training module, configured to train the multi-modal joint audio detection model according to the multi-modal audio feature sample data; Among them, the audio sample data includes first audio sample sub-data and second audio sample sub-data; the first audio sample sub-data and the second audio sample sub-data are audio sample data of different states of the same detection object; the detection object includes a motor; the first audio sample sub-data includes motor forward rotation audio sample sub-data, and the second audio sample sub-data includes motor reverse rotation audio sample sub-data; The multi-modal audio feature extraction module is further configured to: Extract a set number of single-modal audio features from the first audio sample sub-data through the multi-modal joint audio detection model; Extract the set number of single-modal audio features from the second audio sample sub-data through the multi-modal joint audio detection model; Concatenate the single-modal audio features of the first audio sample sub-data and the single-modal audio features of the second audio sample sub-data to obtain concatenated single-modal audio feature sample data; among them, the concatenated single-modal audio feature sample data is feature data obtained by concatenating and fusing the single-modal audio features of a certain same dimension of the first audio sample sub-data and the second audio sample sub-data; Constitute the multi-modal audio feature sample data according to each of the concatenated single-modal audio feature sample data.
11. The apparatus according to claim 10, wherein, The audio sample data acquisition module is further configured to: Acquire original audio data; among them, each of the original audio data includes at least two types of original audio sub-data; Perform data preprocessing on the original audio data to obtain preprocessed audio data; Acquire the audio sample label of the preprocessed audio data; Mark the preprocessed audio data according to the audio sample label to obtain the audio sample data.
12. The device according to claim 11, wherein, The audio sample data acquisition module is further configured to: Perform normalization processing on the original audio data to obtain normalized audio data; Perform truncation processing on the normalized audio data according to the noise interference data in the normalized audio data to obtain first segmented audio data; Delete the offset audio data in the first segmented audio data to obtain second segmented audio data; Perform standardization processing on the data length of the second segmented audio data to obtain the preprocessed audio data.
13. The apparatus according to claim 12, wherein The audio sample data acquisition module is further configured to: When it is determined that the data length of the second segmented audio data is less than the segmented audio standard length, determine the padding length of the second segmented audio data according to the data length of the second segmented audio data and the segmented audio standard length; Collect the collected audio data of the padding length from the second segmented audio data; And pad the second segmented audio data according to the collected audio data to obtain the preprocessed audio data; Or When it is determined that the data length of the second segmented audio data is greater than the segmented audio standard length, intercept the segmented audio standard length of intercepted audio data from the second segmented audio data; and use the intercepted audio data as the preprocessed audio data.
14. The device according to claim 10, wherein, The unimodal audio feature includes at least one of the following: time-frequency map feature, Mel cepstral coefficient MFCC feature, Mel MEL filter bank feature, time-domain mathematical statistical feature, frequency-domain mathematical statistical feature, and short-time frame feature.
15. The apparatus according to claim 14, wherein The time-domain mathematical statistical feature includes at least one of the following: mean, peak value, rectified average value, root mean square, kurtosis, skewness, kurtosis, waveform factor, impulse factor, root mean square amplitude, margin factor, and peak factor; and / or The frequency-domain mathematical statistical feature includes at least one of the following: total energy, frequency centroid, frequency spectrum line, root mean square frequency, kurtosis, skewness, kurtosis, waveform factor, impulse factor, root mean square amplitude, margin factor, and peak factor; and / or The short-time frame feature includes short-time energy, short-time average amplitude, short-time average zero-crossing rate, short-time autocorrelation, and short-time average amplitude difference.
16. The device according to any one of claims 10 and 14 - 15, wherein, The multimodal joint audio detection model includes a result summary model and a set number of audio detection models; the multimodal joint audio detection model training module is further configured to: Input each spliced unimodal audio feature sample data of the multimodal audio feature sample data into each audio detection model of the multimodal joint audio detection model to obtain the unimodal audio detection results of each audio detection model; Splice the unimodal audio detection results of each audio detection model to obtain a spliced unimodal audio detection result; Input the spliced unimodal audio detection result into the result summary model of the multimodal joint audio detection model to obtain a target audio detection result; Calculate the model loss according to the target audio detection result, and update the model parameters of the multimodal joint audio detection model according to the model loss.
17. The apparatus according to claim 16, wherein, The multimodal joint audio detection model training module is further configured to: Determine the number of first audio detection results and the number of second audio detection results through the result summary model according to the spliced unimodal audio detection result; Determine the target audio detection result through the result summary model according to the number of first audio detection results and the number of second audio detection results, and the weights of each spliced unimodal audio detection result.
18. An audio detection device, comprising: An audio data to be detected acquisition module, configured to acquire audio data to be detected; wherein each of the audio data to be detected includes at least two audio sub-data to be detected; An audio detection result acquisition module, configured to input the audio data to be detected into a multi-modal joint audio detection model for audio detection to obtain an audio detection result of the audio data to be detected; Wherein, the multi-modal joint audio detection model is trained by the audio detection model training device according to any one of claims 10-17.
19. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the audio detection model training method according to any one of claims 1-8 or the audio detection method according to claim 9.
20. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the audio detection model training method according to any one of claims 1-8 or the audio detection method according to claim 9.
21. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the audio detection model training method according to any one of claims 1-8 or the audio detection method according to claim 9 is implemented.
Citation Information
Patent Citations
Audio recognition model training method, audio recognition method and device, and equipment
CN111508480A