A method and system for constructing a voice classification target model

By constructing local feature processing of multi-resolution sound images and auxiliary modal spatiotemporal matrix, dynamic sound index is generated, and the problem of insufficient accuracy of traditional sound classification methods in complex environments is solved, and efficient adaptability and accuracy of sound classification is achieved.

CN120067775BActive Publication Date: 2025-07-22HANGZHOU LIFANG CULTURE MEDIA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510541912.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-22
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Traditional sound classification methods cannot effectively capture the time dependence of sound, and existing multimodal methods cannot dynamically adapt to the importance changes of each modal in complex environments. The model is susceptible to noise, reverb, and sensor differences in real environments, reducing generalization capabilities.

Method used

By obtaining original sound data and auxiliary modal data, multi-resolution sound images and auxiliary modal spatiotemporal matrix processing are performed, combined with local feature extraction and hierarchical feature sequence space construction, dynamic sound index is generated, and a sound classification target model is constructed.

Benefits of technology

It improves the accuracy of sound classification, reduces the influence of factors such as noise, reverberation, sensor differences in real environments, and enhances the model's adaptability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067775B_ABST
    Figure CN120067775B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for constructing a sound classification target model, which relates to the technical field of sound signal processing, and includes: obtaining original sound data and auxiliary modality data, performing sound data conversion processing on the original sound data to obtain corresponding multi-resolution sound image data; processing the auxiliary modality data to obtain corresponding auxiliary modality spatio-temporal matrices; performing local feature processing on the multi-resolution sound image data and the auxiliary modality spatio-temporal matrices to obtain local sound image features and auxiliary modality features; constructing a hierarchical feature sequence space according to the local sound image features and the auxiliary modality features to obtain corresponding fused feature sequences; constructing a sound classification target model according to the fused feature sequences, generating a dynamic sound index according to the sound classification target model, and further obtaining a sound classification result, thereby improving the accuracy of sound classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sound signal processing, and specifically to a method and system for constructing a sound classification target model. Background Art

[0002] After converting sound into a static spectrogram by traditional methods, an image classification model (such as CNN) is directly used for processing, resulting in the compression of temporal dynamic features (such as the prosodic changes of speech and the continuous evolution of environmental sounds) into two-dimensional spatial information, making it difficult to capture the time dependence of sound. For example, the time-domain transient characteristics of sudden noise are manifested as isolated high-frequency points in the spectrogram, and traditional CNNs may misclassify them as background textures.

[0003] Existing multimodal methods (such as early fusion, late fusion) usually adopt fixed weights or simple splicing strategies and cannot dynamically adapt to the importance changes of each modality in complex environments. For example, in a high-noise scenario, the data of vibration sensors may be more reliable than audio data, but traditional methods cannot automatically increase their weights.

[0004] After the model is trained under laboratory conditions and deployed to a real environment (such as the ocean, industrial sites), it is vulnerable to factors such as noise, reverberation, and sensor differences. For example, the difference in sound speed caused by temperature and salinity changes in underwater sonar signals will significantly reduce the generalization ability of the model.

[0005] Therefore, a method and system for constructing a sound classification target model are provided. Summary of the Invention

[0006] In order to solve the above technical problems, the purpose of the present invention is to provide a method and system for constructing a sound classification target model.

[0007] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing a sound classification target model, comprising:

[0008] Obtain the original sound data and auxiliary modality data, perform sound data conversion processing on the original sound data to obtain corresponding multi-resolution sound image data; process the auxiliary modality data to obtain a corresponding auxiliary modality spatio-temporal matrix;

[0009] Perform local feature processing on the multi-resolution sound image data and the auxiliary modality spatio-temporal matrix to obtain local sound image features and auxiliary modality features;

[0010] Construct a hierarchical feature sequence space according to the local sound image features and the auxiliary modality features to obtain a corresponding fusion feature sequence;

[0011] Construct a sound classification target model according to the fusion feature sequence, generate a dynamic sound index according to the sound classification target model, and then obtain a sound classification result.

[0012] According to one preferred embodiment of the present invention, the process of obtaining the original sound data and the auxiliary modality data includes:

[0013] Preset a data acquisition device, which is composed of several data acquisition units, and set an acquisition period. The data acquisition units include an original sound data acquisition unit, a device vibration data acquisition unit, and an environmental temperature data acquisition unit;

[0014] The original sound data acquisition unit is used to acquire the original sound data;

[0015] The device vibration data acquisition unit is used to acquire the sensor vibration sound data when the data acquisition unit is working;

[0016] The environmental temperature data acquisition unit is used to acquire the environmental temperature within the current acquisition period;

[0017] Record the sensor vibration sound data and the environmental temperature as auxiliary modality data.

[0018] According to one preferred embodiment of the present invention, the process of performing sound data conversion processing on the original sound data includes:

[0019] Obtain the original sound data;

[0020] Perform preprocessing on the original sound data. Based on the Hamming window function, divide the preprocessed original sound data into multiple short time segments, and perform Fourier transform on each segment to obtain a time-frequency representation , the time-frequency representation is:

[0021] ; where is the original sound data, is the Hamming window function, is the frame shift, is the window length, m is the frame index, and k is the frequency index;

[0022] Based on the Griffin-Lim algorithm, obtain a high-resolution spectrogram with higher accuracy;

[0023] Based on the Wasserstein GAN technology, optimize the high-resolution spectrogram to improve the quality of the high-resolution spectrogram, and record it as multi-resolution sound image data.

[0024] According to one preferred embodiment of the present invention, the process of processing the auxiliary modal data includes:

[0025] Preset several feature extraction times;

[0026] Based on the filtering processing technology, remove the noise interference of the sensor vibration sound data and retain the useful signal components. Based on the wavelet denoising technology, improve the quality of the filtered sensor vibration sound data;

[0027] Extract the amplitude features, frequency features, and energy features from the sensor vibration sound data after filtering and denoising processing. Synthesize matrices according to the amplitude features, frequency features, and energy features extracted at different feature extraction times to form the spatio-temporal matrix of the sensor vibration sound data;

[0028] Preset Feature extraction is performed on the sensor vibration sound data at time points, and

[0029] features are extracted at each time point to generate the spatio-temporal matrix of the sensor vibration sound data;

[0030] Arrange the environmental temperature in the order of the feature extraction time to form a one-dimensional vector;

[0031] According to one preferred embodiment of the present invention, the process of performing local feature processing on the multi-resolution sound image data and the auxiliary modal spatio-temporal matrix includes:

[0032] Based on the gray-level co-occurrence matrix technology, generate the local sound image feature set of the multi-resolution sound image data, and obtain the texture feature similarity between the local regions of each pixel point and the local regions of other adjacent pixel points according to the local sound image feature set;

[0033] Perform a linear mapping on the auxiliary modal spatio-temporal matrix. The process is as follows:

[0034] Obtain the feature dimension ;

[0035] Preset the target feature dimension ;

[0036] Map the feature dimension of the auxiliary modal spatio-temporal matrix to the target dimension , and the calculation process of the linear mapping is as follows:

[0037] ; where is the spatio-temporal matrix, is the weight matrix, with a shape of , b is the bias vector, with a shape of , is the output matrix after mapping, with a shape of , is the auxiliary modal spatio-temporal matrix of the batch size, is the auxiliary modal spatio-temporal matrix of the sequence length;

[0038] Based on the Transformer encoder, and according to the output matrix , output the auxiliary modal features.

[0039] According to one preferred embodiment of the present invention, the process of constructing the hierarchical feature sequence space includes:

[0040] Obtain the local sound image feature set and the auxiliary modal features;

[0041] Fuse the local sound image feature set and the auxiliary modal features to obtain a fused feature set, denoted as , and the fused feature set is:

[0042] ; where is the local sound image feature set, is the auxiliary modal feature, , are the weight coefficients;

[0043] Based on the small neural network technology, dynamically allocate the weight coefficients of the local sound image feature set and the auxiliary modal features;

[0044] According to the fused feature set and the dynamically allocated weight coefficients, generate the hierarchical feature sequence space, and then the corresponding fused feature sequence.

[0045] According to one preferred embodiment of the present invention, the process of constructing the sound classification target model includes:

[0046] Obtain several groups of fused feature sequences;

[0047] Preset the standard fused feature sequence;

[0048] Group and label several groups of fused feature sequences, denoted as is a natural number;

[0049] Take groups of several groups of fused feature sequences as sample data, and is less than natural numbers, and use the sample data to obtain the mean of the sample data, denoted as the sample set;

[0050] Use the remaining several groups of fusion feature sequences and the standard fusion feature sequence as the test set;

[0051] According to the sample set and the test set, form a training sample set;

[0052] Based on the convolutional neural network, construct a standard sound classification model;

[0053] And input the training sample set into the standard sound classification model to train the standard sound classification model, and obtain the trained standard sound classification model, and denote the trained standard sound classification model as the sound classification target model.

[0054] According to one preferred embodiment of the present invention, the process of generating a dynamic sound index according to the sound classification target model and then obtaining the sound classification result includes:

[0055] Generate a dynamic sound index under the current environmental factor conditions according to the sound classification target model , the dynamic sound index is:

[0056] ; where represents the actual environmental temperature, represents the standard environmental temperature, represents the proportionality coefficient of the temperature influence degree, represents the maximum amplitude value of the sensor vibration sound data, represents the texture feature similarity between the local area of the i-th pixel point and the local area of the j-th pixel point, represents the current standard environmental factor, represents the actual environmental factor;

[0057] Preset the standard dynamic sound index ;

[0058] If the dynamic sound index , then the sound under the current environmental factor conditions meets the classification standard;

[0059] If the dynamic sound index , then the sound under the current environmental factor conditions does not meet the classification standard.

[0060] The second aspect of the present invention also provides a sound classification target model construction system, which executes the above-mentioned method for constructing a sound classification target model, including: a sound data acquisition module, a sound data processing module, a sound data feature processing module, a sound data storage module, and a classification target model construction module;

[0061] The sound data acquisition module is used to acquire the original sound data and auxiliary modality data;

[0062] The sound data processing module is used to process the original sound data and auxiliary modality data to obtain multi - resolution sound image data and auxiliary modality spatio - temporal matrices;

[0063] The sound data feature processing module is used to perform local feature processing on the multi - resolution sound image data and auxiliary modality spatio - temporal matrices to obtain local sound image features and auxiliary modality features, and construct a hierarchical feature sequence space to obtain a fused feature sequence;

[0064] The sound data storage module is used to store the fused feature sequence;

[0065] The classification target model construction module is used to construct a sound classification target model according to the fused feature sequence, generate a dynamic sound index according to the sound classification target model, and then obtain a sound classification result.

[0066] The present invention further provides a computer - readable storage medium, which stores a computer program, and the computer program can be executed by a processor to implement the above - mentioned method for constructing a sound classification target model.

[0067] Compared with the prior art, the beneficial effects of the present invention are as follows: acquiring the original sound data and auxiliary modality data, performing sound data conversion processing on the original sound data to obtain corresponding multi - resolution sound image data; processing the auxiliary modality data to obtain corresponding auxiliary modality spatio - temporal matrices; dynamically adapting to the importance changes of each modality in a complex environment, and improving the sound recognition ability.

[0068] Performing local feature processing on the multi - resolution sound image data and auxiliary modality spatio - temporal matrices to obtain local sound image features and auxiliary modality features; constructing a hierarchical feature sequence space according to the local sound image features and auxiliary modality features to obtain a corresponding fused feature sequence; constructing a sound classification target model according to the fused feature sequence, generating a dynamic sound index according to the sound classification target model, and then obtaining a sound classification result. After training under laboratory conditions and being deployed to a real environment (such as the ocean, industrial site), it is less affected by factors such as noise, reverberation, and sensor differences, and improves the accuracy of sound classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.

[0070] Figure 1 It is a schematic diagram of a method for constructing a sound classification target model.

[0071] Figure 2 It is a module diagram of a sound classification target model construction system. Detailed implementation manners

[0072] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations. The basic principles defined in the following description of the present invention can be applied to other implementation manners, variations, improvements, equivalent manners, and other technical solutions that do not depart from the spirit and scope of the present invention.

[0073] It can be understood that the term "one" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, and in other embodiments, the number of this element can be multiple. The term "one" cannot be understood as a limitation on the number.

[0074] As Figure 1 shown, a method for constructing a sound classification target model includes the following steps:

[0075] Obtain the original sound data and auxiliary modality data, perform sound data conversion processing on the original sound data to obtain corresponding multi-resolution sound image data; process the auxiliary modality data to obtain corresponding auxiliary modality spatio-temporal matrices;

[0076] Perform local feature processing on the multi-resolution sound image data and the auxiliary modality spatio-temporal matrices to obtain local sound image features and auxiliary modality features;

[0077] Construct a hierarchical feature sequence space according to the local sound image features and the auxiliary modality features to obtain corresponding fused feature sequences;

[0078] Construct a sound classification target model according to the fused feature sequences, generate a dynamic sound index according to the sound classification target model, and further obtain a sound classification result.

[0079] It should be further noted that, in the specific implementation process, the specific process of obtaining the original sound data and the auxiliary modality data includes:

[0080] A preset data acquisition device, which is composed of several data acquisition units and is set with an acquisition period. The data acquisition units include an original sound data acquisition unit, a device vibration data acquisition unit, and an environmental temperature data acquisition unit;

[0081] The original sound data acquisition unit is used to acquire original sound data; the device vibration data acquisition unit is used to acquire sensor vibration sound data when the data acquisition unit is working; the environmental temperature data acquisition unit is used to acquire the environmental temperature within the current acquisition period. It should be further noted that during the process of acquiring sound data by the data acquisition device, the corresponding data acquisition unit will generate subtle vibration sound data; the sensor vibration sound data and the environmental temperature are recorded as auxiliary modal data.

[0082] It should be further noted that in the specific implementation process, the specific process of performing sound data conversion processing on the original sound data includes:

[0083] Obtain the original sound data;

[0084] Perform preprocessing on the original sound data. The specific process includes:

[0085] Select a sampling frequency of 44100Hz and a quantization bit number of 24 bits to sample and quantize the original sound data. It should be further noted that select appropriate sampling frequency and quantization bit number according to actual needs;

[0086] Based on a filter, remove the noise interference in the original sound data after sampling and quantization, and retain the frequency components related to sound classification;

[0087] Normalize the amplitude of the original sound data after removing noise interference to [-1, 1] to avoid affecting subsequent processing due to too large signal amplitude differences;

[0088] Based on the Hamming window function, divide the preprocessed original sound data into multiple short time segments, and perform Fourier transform on each segment to obtain the time-frequency representation , the time-frequency representation is:

[0089] ; where is the original sound data, is the Hamming window function, is the frame shift, is the window length, m is the frame index, and k is the frequency index;

[0090] Adjust the window length and frame shift of the Hamming window function to generate Mel spectrograms with different resolutions, where the Mel spectrograms include those with resolutions of 256×256 and 128×128. It should be further noted that Mel spectrograms with different resolutions can capture features of different scales of the audio signal. High-resolution spectrograms can retain more detailed information, while low-resolution spectrograms can highlight the overall feature trends.

[0091] Estimate the phase information from the time-frequency representation initially, randomly initialize the phase, and then synthesize a signal from the time-frequency representation and the current initialized phase, and then perform STFT calculation on the synthesized signal to obtain a new time-frequency representation and phase, continuously iterate and update the phase until the convergence condition is met; obtain a high-resolution spectrogram with higher accuracy.

[0092] Optimize the high-resolution spectrogram based on the Wasserstein GAN technology. By minimizing the Wasserstein distance, improve the quality of the high-resolution spectrogram, and denote it as multi-resolution sound image data.

[0093] It should be further noted that in the specific implementation process, the specific process of processing the auxiliary modal data includes:

[0094] Preset several feature extraction times.

[0095] Based on the filtering processing technology, remove the noise interference of the sensor vibration sound data and retain the useful signal components. Based on the wavelet denoising technology, improve the quality of the filtered sensor vibration sound data.

[0096] Extract features from the sensor vibration sound data after filtering and denoising, specifically including:

[0097] 1. Amplitude feature extraction, obtain the maximum value, mean value, and root mean square value of the sensor vibration sound data within the feature extraction time.

[0098] 2. Frequency feature extraction, based on the Fourier transform, convert the sensor vibration sound data within the feature extraction time to the frequency domain, and calculate the main frequency and spectral energy distribution of the sensor vibration sound data.

[0099] 3. Energy feature extraction, perform wavelet packet decomposition on the sensor vibration sound data within the feature extraction time, and calculate the energy of each wavelet packet node.

[0100] Extract the amplitude feature, frequency feature, and energy feature of the synthesized matrix according to the time of different features to form the spatio-temporal matrix of the sensor vibration sound data;

[0101] Preset Feature extraction is performed on the sensor vibration sound data at time points, and features are extracted at each time point to generate the spatio-temporal matrix of the sensor vibration sound data which is: ;

[0102] Arrange the environmental temperature in the order of the feature extraction time to form a one-dimensional vector to ensure the correct time order of the data for matching with the sensor vibration sound data;

[0103] According to the spatio-temporal matrix and the one-dimensional vector generate the corresponding auxiliary modal spatio-temporal matrix where , are weight coefficients.

[0104] It should be further noted that in the specific implementation process, the specific process of performing local feature processing on the multi-resolution sound image data and the auxiliary modal spatio-temporal matrix includes:

[0105] The process of generating the local sound image feature set of the multi-resolution sound image data based on the gray-level co-occurrence matrix technology includes:

[0106] Calculate the gray-level co-occurrence matrix around each pixel point of the multi-resolution sound image data by means of a sliding window. The specific details are as follows: Quantify the gray level of the multi-resolution sound image data into discrete gray levels, divide the gray value range into several levels, for example, an 8-bit image can be divided into 16, 32, 64 levels, and define the parameters required for the gray-level co-occurrence matrix, including distance and direction, which are determined according to the actual application requirements. The specific steps will not be elaborated here. Subsequently, based on the calculated gray-level co-occurrence matrix, extract a series of local sound image features to construct a local sound image feature set. The local sound image features in the local sound image feature set include but are not limited to:

[0107] Contrast: A statistical feature that describes the contrast of pixels with different gray levels in an image, a feature that measures the roughness of the image texture, reflects the clarity of the image and the depth of the texture grooves. The deeper the texture grooves, the greater the contrast and the clearer the visual effect. On the contrary, if the contrast is small, the grooves are shallow and the effect is blurred;

[0108] Correlation: Describes the degree of correlation of pixels with different gray levels in an image. It measures the similarity of the elements of the spatial gray-level co-occurrence matrix in the row or column direction. Therefore, the magnitude of the correlation value reflects the local gray-level correlation in the image. When the matrix element values are uniformly equal, the correlation value is large. On the contrary, if the matrix pixel values vary greatly, the correlation value is small. If there are horizontal textures in the image, the correlation of the horizontal matrix is greater than that of the other matrices;

[0109] Energy: Describes the degree of uniformity of the pixel gray-level distribution in an image, measures the randomness contained in the image, and represents the complexity of the image. When all the values of the co-occurrence matrix are equal or the pixel values show the maximum randomness, the entropy is the largest;

[0110] Homogeneity: Describes the similarity of the gray levels of adjacent pixels in an image; reflects the homogeneity of the image texture, measures the amount of local change in the image texture. A large value indicates that there is little change between different regions of the image texture and it is very uniform locally;

[0111] Entropy: Describes the degree of uncertainty of the image texture, measures the randomness contained in the image, and represents the complexity of the image. When all the values of the co-occurrence matrix are equal or the pixel values show the maximum randomness, the entropy is the largest;

[0112] Inverse variance: Reflects the clarity and regularity of the texture, with clear texture and strong regularity;

[0113] It should be further noted that in the specific implementation process, the calculation formula for obtaining the texture feature similarity between the local regions of each pixel point and the local regions of other adjacent pixel points is:

[0114] ;

[0115] Where, represents the texture feature similarity between the local region of the i-th pixel point and the local region of the j-th pixel point, represents the contrast between the local region of the i-th pixel point and the local region of the j-th pixel point, represents the energy of the local region of the i-th pixel point, represents the energy of the local region of the j-th pixel point, represents the entropy of the local region of the i-th pixel point, represents the entropy of the local region of the j-th pixel point, represents the inverse variance of the local region of the i-th pixel point, represents the inverse variance of the local region of the j-th pixel point, 、 、 、 and represent weight factors.

[0116] Perform a linear mapping on the auxiliary modal spatio-temporal matrix, and the specific process is as follows:

[0117] Obtain the characteristic dimension of the auxiliary modal spatio-temporal matrix ; It should be further noted that the values of the characteristic dimension are 128, 256, 512, etc., and are selected according to actual needs;

[0118] Preset the target characteristic dimension ;

[0119] Map the characteristic dimension of the auxiliary modal spatio-temporal matrix to the target dimension , and the calculation process of the linear mapping is as follows:

[0120] ; Among them, is the weight matrix, with a shape of , b is the bias vector, with a shape of , is the output matrix after mapping, with a shape of , is the auxiliary modal spatio-temporal matrix 's batch size, is the auxiliary modal spatio-temporal matrix 's sequence length;

[0121] Based on the Transformer encoder, learn the output matrix and output the auxiliary modal features.

[0122] It should be further noted that in the specific implementation process, the specific process of constructing the hierarchical feature sequence space includes:

[0123] Obtain the local sound image feature set and the auxiliary modal features;

[0124] Fuse the local sound image feature set and the auxiliary modal features to obtain a fused feature set, denoted as , and the fused feature set is:

[0125] ; Among them, is the local sound image feature set, is the auxiliary modal feature, , are the weight coefficients;

[0126] Based on small neural network technology, the weight coefficients of the local sound image feature set and the auxiliary modal features are dynamically allocated. It should be further noted that the specific allocation of the weight coefficients can be modified according to the actual specific situation;

[0127] According to the fusion feature set and the dynamically allocated weight coefficients, a hierarchical feature sequence space is generated, and then the corresponding fusion feature sequence.

[0128] It should be further noted that in the specific implementation process, the specific process of constructing the sound classification target model includes:

[0129] Obtain several groups of fusion feature sequences;

[0130] Preset a standard fusion feature sequence; it should be further noted that the standard fusion feature sequence includes the standard fusion feature sequences of environmental factors such as noise and reverberation;

[0131] Group and label several groups of fusion feature sequences, denoted as is a natural number;

[0132] Take groups of several groups of fusion feature sequences as sample data, and is a natural number less than , and use the sample data to obtain the sample data mean, denoted as the sample set;

[0133] Take the remaining several groups of fusion feature sequences and the standard fusion feature sequence as the test set;

[0134] According to the sample set and the test set, form a training sample set;

[0135] Based on the convolutional neural network, construct a standard sound classification model;

[0136] And input the training sample set into the standard sound classification model, train the standard sound classification model, obtain the trained standard sound classification model, and denote the trained standard sound classification model as the sound classification target model.

[0137] It should be further noted that in the specific implementation process, according to the sound classification target model, the specific process of generating the dynamic sound index and then obtaining the sound classification result includes:

[0138] According to the sound classification target model, generate the dynamic sound index under the current environmental factor conditions , the dynamic sound index is:

[0139] ; where, represents the actual ambient temperature, represents the standard ambient temperature, represents the proportionality coefficient of the temperature influence degree, represents the maximum amplitude value of the sensor vibration sound data, represents the texture feature similarity between the local area of the i-th pixel point and the local area of the j-th pixel point, represents the current standard ambient factor, represents the actual ambient factor;

[0140] Preset standard dynamic sound index ;

[0141] If the dynamic sound index , then the sound under the current ambient factor conditions conforms to the classification standard;

[0142] If the dynamic sound index , then the sound under the current ambient factor conditions does not conform to the classification standard.

[0143] As Figure 2 shown, a sound classification target model construction system includes: a sound data acquisition module, a sound data processing module, a sound data feature processing module, a sound data storage module, and a classification target model construction module;

[0144] The sound data acquisition module is used to acquire the original sound data and auxiliary modal data;

[0145] The sound data processing module is used to process the original sound data and auxiliary modal data to obtain multi-resolution sound image data and auxiliary modal spatio-temporal matrices;

[0146] The sound data feature processing module is used to perform local feature processing on the multi-resolution sound image data and auxiliary modal spatio-temporal matrices to obtain local sound image features and auxiliary modal features, and construct a hierarchical feature sequence space to obtain a fused feature sequence;

[0147] The sound data storage module is used to store the fused feature sequence;

[0148] The classification target model construction module is used to construct a sound classification target model according to the fused feature sequence, generate a dynamic sound index according to the sound classification target model, and further obtain a sound classification result.

[0149] Embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. Embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the methods of the present application are performed. It should be noted that the computer-readable medium in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wire segments, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire segments, optical cables, RF, etc., or any suitable combination of the above.

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0151] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and illustrated in the embodiments. Without departing from the said principles, the embodiments of the present invention may have any variations or modifications.

Claims

1. A method for constructing a voice classification target model, characterized in that Including: Obtain the original sound data and auxiliary modality data, perform sound data conversion processing on the original sound data to obtain corresponding multi-resolution sound image data; Process the auxiliary modality data to obtain a corresponding auxiliary modality spatio-temporal matrix; Perform local feature processing on the multi-resolution sound image data and the auxiliary modality spatio-temporal matrix to obtain local sound image features and auxiliary modality features; Fuse the local sound image feature set and the auxiliary modality features to obtain a fused feature set, denoted as , the fused feature set is as follows: ; wherein, is a local sound image feature set, is an auxiliary modality feature, , are weight coefficients; Based on small neural network technology, dynamically allocate the weight coefficients of the local sound image feature set and the auxiliary modality features; According to the described fusion feature set and the dynamically allocated weight coefficients, a hierarchical feature sequence space is generated, and then the corresponding fusion feature sequence; According to the fusion feature sequence, construct a sound classification target model, and according to the sound classification target model, generate a dynamic sound index, and then obtain a sound classification result.

2. The method for constructing a voice classification target model according to claim 1, wherein The process of obtaining the original sound data and the auxiliary modality data includes: Preset a data acquisition device, which is composed of several data acquisition units, and set an acquisition period. The data acquisition unit includes an original sound data acquisition unit, a device vibration data acquisition unit, and an environmental temperature data acquisition unit; The original sound data acquisition unit is used to acquire the original sound data; The device vibration data acquisition unit is used to acquire the sensor vibration sound data when the data acquisition unit is working; The environmental temperature data acquisition unit is used to acquire the environmental temperature during the current acquisition period; Record the sensor vibration sound data and the environmental temperature as auxiliary modality data.

3. The method for constructing a sound classification target model according to claim 2, wherein The process of performing sound data conversion processing on the original sound data includes: Obtain the original sound data; Preprocess the original sound data, based on the Hamming window function, divide the preprocessed original sound data into multiple short time segments, perform Fourier transform on each segment to obtain the time-frequency representation , the time-frequency representation is: ; wherein, is the original audio data, is the Hamming window function, is the frame shift, is the window length, m is the frame index, and k is the frequency index; Based on the Griffin-Lim algorithm, obtain a high-resolution spectrogram with higher accuracy; Based on the Wasserstein GAN technology, optimize the high-resolution spectrogram, improve the quality of the high-resolution spectrogram, and record it as multi-resolution sound image data.

4. The method for constructing a voice classification target model according to claim 3, wherein The process of processing the auxiliary modality data includes: Preset several feature extraction times; Based on the filtering processing technology, remove the noise interference of the sensor vibration sound data, retain the useful signal components, and based on the wavelet denoising technology, improve the quality of the filtered sensor vibration sound data; Extract amplitude features, frequency features, and energy features from the sensor vibration sound data after filtering and denoising processing, and synthesize matrices according to the amplitude features, frequency features, and energy features extracted at different feature extraction times to form a spatio-temporal matrix of the sensor vibration sound data; Preset Feature extraction was performed on the sensor vibration sound data at time points, and features were extracted at each time point to generate a spatio-temporal matrix of the sensor vibration sound data; Arrange the environmental temperature in the order of the feature extraction times to form a one-dimensional vector; Generate a corresponding auxiliary modality spatio-temporal matrix according to the spatio-temporal matrix and the one-dimensional vector.

5. A method for constructing a voice classification target model according to claim 4, characterized in that The process of performing local feature processing on the multi-resolution sound image data and the auxiliary modality spatio-temporal matrix includes: Based on the gray-level co-occurrence matrix technology, generate a local sound image feature set of the multi-resolution sound image data, and obtain the texture feature similarity between the local regions of each pixel point and the local regions of other adjacent pixel points according to the local sound image feature set; The process of performing linear mapping on the auxiliary modality spatio-temporal matrix is: Obtain the characteristic dimension of the auxiliary modal spatio-temporal matrix ; Preset target feature dimension ; Map the feature dimension of the auxiliary modal spatio-temporal matrix to the target dimension , and the calculation process of the linear mapping is as follows: ; among them, is a spatio-temporal matrix, is a weight matrix, with a shape of , b is a bias vector, with a shape of , is the output matrix after mapping, with a shape of , is the auxiliary modality spatio-temporal matrix 's batch size, is the auxiliary modality spatio-temporal matrix 's sequence length; Based on the Transformer encoder and according to the output matrix , the auxiliary modal features are output.

6. The method for constructing a voice classification target model according to claim 5, wherein The process of constructing a sound classification target model includes: Obtain several groups of fusion feature sequences; Preset a standard fusion feature sequence; Group and label several sets of fusion feature sequences, denoted as is a natural number; Group several groups of fusion feature sequences as sample data, and is a natural number less than , and use the sample data to obtain the sample data mean, denoted as the sample set;​ Use the remaining several groups of fusion feature sequences and the standard fusion feature sequence as the test set; According to the sample set and the test set, form a training sample set; Based on the convolutional neural network, construct a standard sound classification model; And input the training sample set into the standard sound classification model to train the standard sound classification model, obtain the trained standard sound classification model, and denote the trained standard sound classification model as the sound classification target model.

7. A method for constructing a voice classification target model according to claim 6, characterized in that The process of generating a dynamic sound index according to the sound classification target model and then obtaining the sound classification result includes: Generate a dynamic sound index under the current environmental factor conditions according to the described sound classification target model , the dynamic sound index is as follows: ; wherein, represents the actual ambient temperature, represents the standard ambient temperature, represents the proportionality coefficient of the temperature influence degree, represents the maximum amplitude value of the sensor vibration sound data, represents the texture feature similarity between the local area of the i-th pixel and the local area of the j-th pixel, represents the current standard environmental factor, represents the actual environmental factor; Preset standard dynamic sound index ; If the dynamic sound index , the sound under the current environmental factor conditions meets the classification standard; If the dynamic sound index , the sound under the current environmental factor conditions does not meet the classification standard.

8. A system for constructing a sound classification target model, the system executes the method for constructing a sound classification target model according to any one of the above claims 1-7, including: A sound data acquisition module, a sound data processing module, a sound data feature processing module, a sound data storage module, and a classification target model construction module; The sound data acquisition module is used to acquire the original sound data and the auxiliary modality data; The sound data processing module is used to process the original sound data and the auxiliary modality data to obtain multi-resolution sound image data and an auxiliary modality spatio-temporal matrix; The voice data feature processing module is used to perform local feature processing on the multi-resolution voice image data and the auxiliary modality spatio-temporal matrix to obtain local voice image features and auxiliary modality features, and fuse the local voice image feature set and the auxiliary modality features to obtain a fused feature set, denoted as , the fused feature set is as follows: ; wherein, is the local sound image feature set, is the auxiliary modality feature, , are the weight coefficients; Based on the small neural network technology, dynamically allocate the weight coefficients of the local sound image feature set and the auxiliary modality features; According to the described fusion feature set and the dynamically allocated weight coefficients, a hierarchical feature sequence space is generated, and then the corresponding fusion feature sequence; The sound data storage module is used to store the fusion feature sequence; The classification target model construction module is used to construct a sound classification target model according to the fusion feature sequence, generate a dynamic sound index according to the sound classification target model, and then obtain the sound classification result.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement a method for constructing a sound classification target model according to any one of claims 1-7 above.

Citation Information

Patent Citations

  • Voice-based video action classification method and related equipment

    CN114529846A

  • Structural damage identification method and system in combination with multi-modal information and artificial intelligence

    CN118552795A