Sound classification target model construction method and system
By constructing a hierarchical feature sequence space and dynamically adjusting feature weights, the problem of difficult sound time dependence and poor environmental adaptability in the prior art is solved, and higher sound classification accuracy and generalization ability are achieved.
Patent Information
- Application Number
- CN202510541912.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The prior art is difficult to effectively capture the time dependence of sound, and it is unable to dynamically adapt to the importance changes of each mode in complex environments. It is susceptible to noise, reverb and sensor differences, reducing the accuracy of sound classification.
By obtaining the original sound data and auxiliary modal data, multi-resolution sound image data and auxiliary modal spatiotemporal matrix are processed, hierarchical feature sequence space is constructed, feature weights are dynamically adjusted, and dynamic sound index is generated to realize sound classification.
It improves the recognition and classification accuracy of sound, reduces the impact of noise and sensor differences in complex environments, and enhances the generalization ability of the model.
Smart Images

Figure CN120067775A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sound signal processing, and specifically to a method and system for constructing a sound classification target model. Background Art
[0002] After converting sound into a static spectrogram by traditional methods, an image classification model (such as CNN) is directly used for processing, resulting in the compression of temporal dynamic features (such as the prosodic changes of speech and the continuous evolution of environmental sounds) into two-dimensional spatial information, making it difficult to capture the time dependence of sound. For example, the time-domain transient characteristics of sudden noise are manifested as isolated high-frequency points in the spectrogram, and traditional CNNs may misclassify them as background textures.
[0003] Existing multimodal methods (such as early fusion, late fusion) usually adopt fixed weights or simple splicing strategies and cannot dynamically adapt to the importance changes of each modality in complex environments. For example, in a high-noise scenario, the data of vibration sensors may be more reliable than audio data, but traditional methods cannot automatically increase their weights.
[0004] After the model is trained under laboratory conditions and deployed to real environments (such as the ocean, industrial sites), it is vulnerable to factors such as noise, reverberation, and sensor differences. For example, the difference in sound speed caused by temperature and salinity changes in underwater sonar signals will significantly reduce the generalization ability of the model.
[0005] Therefore, a method and system for constructing a sound classification target model are provided. Summary of the Invention
[0006] In order to solve the above technical problems, the purpose of the present invention is to provide a method and system for constructing a sound classification target model.
[0007] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing a sound classification target model, comprising: Obtaining original sound data and auxiliary modality data, performing sound data conversion processing on the original sound data to obtain corresponding multi-resolution sound image data; processing the auxiliary modality data to obtain a corresponding auxiliary modality spatio-temporal matrix; Performing local feature processing on the multi-resolution sound image data and the auxiliary modality spatio-temporal matrix to obtain local sound image features and auxiliary modality features; Constructing a hierarchical feature sequence space according to the local sound image features and the auxiliary modality features to obtain a corresponding fused feature sequence; Constructing a sound classification target model according to the fused feature sequence, generating a dynamic sound index according to the sound classification target model, and further obtaining a sound classification result.
[0008] According to one preferred embodiment of the present invention, the process of obtaining the original sound data and the auxiliary modality data includes: Preset a data acquisition device, which is composed of a plurality of data acquisition units, and set an acquisition period. The data acquisition units include an original sound data acquisition unit, a device vibration data acquisition unit, and an environmental temperature data acquisition unit; The original sound data acquisition unit is used to acquire the original sound data; The device vibration data acquisition unit is used to acquire the sensor vibration sound data when the data acquisition unit is working; The environmental temperature data acquisition unit is used to acquire the environmental temperature within the current acquisition period; Record the sensor vibration sound data and the environmental temperature as the auxiliary modality data.
[0009] According to one preferred embodiment of the present invention, the process of performing sound data conversion processing on the original sound data includes: Obtain the original sound data; Perform preprocessing on the original sound data. Based on the Hamming window function, divide the preprocessed original sound data into multiple short time segments, and perform Fourier transform on each segment to obtain the time-frequency representation , the time-frequency representation is: ; where is the original sound data, is the Hamming window function, is the frame shift, is the window length, m is the frame index, and k is the frequency index; Based on the Griffin-Lim algorithm, obtain a high-resolution spectrogram with higher accuracy; Based on the Wasserstein GAN technology, optimize the high-resolution spectrogram to improve the quality of the high-resolution spectrogram, and record it as the multi-resolution sound image data.
[0010] According to one preferred embodiment of the present invention, the process of processing the auxiliary modality data includes: Preset several feature extraction times; Based on the filtering processing technology, remove the noise interference of the sensor vibration sound data and retain the useful signal components. Based on the wavelet denoising technology, improve the quality of the filtered sensor vibration sound data; Perform amplitude feature extraction, frequency feature extraction, and energy feature extraction on the filtered and denoised sensor vibration sound data. Synthesize matrices according to the amplitude feature extraction, frequency feature extraction, and energy feature extraction at different feature extraction times to form the spatio-temporal matrix of the sensor vibration sound data; Preset Feature extraction was performed on the sensor vibration sound data at preset time points, and features were extracted at each time point to generate a spatio-temporal matrix of the sensor vibration sound data; The environmental temperature was arranged in the order of the feature extraction time to form a one-dimensional vector;
[0011] According to one preferred embodiment of the present invention, the process of performing local feature processing on the multi-resolution sound image data and the auxiliary modal spatio-temporal matrix includes: Based on the gray-level co-occurrence matrix technology, a local sound image feature set of the multi-resolution sound image data was generated, and the texture feature similarity between the local regions of each pixel point and the local regions of other adjacent pixel points was obtained according to the local sound image feature set; Perform a linear mapping on the auxiliary modal spatio-temporal matrix. The process is as follows: Obtain the feature dimension of the auxiliary modal spatio-temporal matrix ; Preset the target feature dimension ; Map the feature dimension of the auxiliary modal spatio-temporal matrix to the target dimension , and the calculation process of the linear mapping is as follows: ; where, is the spatio-temporal matrix, is the weight matrix, with a shape of , b is the bias vector, with a shape of , is the output matrix after mapping, with a shape of , is the batch size of the auxiliary modal spatio-temporal matrix , is the sequence length of the auxiliary modal spatio-temporal matrix ; Based on the Transformer encoder, and according to the output matrix , output the auxiliary modal features.
[0012] According to one preferred embodiment of the present invention, the process of constructing the hierarchical feature sequence space includes: Obtain the local sound image feature set and the auxiliary modal features; Fuse the local sound image feature set and the auxiliary modal features to obtain a fused feature set, denoted as , and the fused feature set is: ; wherein, is the local sound image feature set, is the auxiliary modality feature, , are the weight coefficients; Based on the small neural network technology, dynamically allocate the weight coefficients of the local sound image feature set and the auxiliary modality feature; According to the fusion feature set and the dynamically allocated weight coefficients, generate a hierarchical feature sequence space, and then the corresponding fusion feature sequence.
[0013] According to one preferred embodiment of the present invention, the process of constructing the sound classification target model includes: Obtain several groups of fusion feature sequences; Preset the standard fusion feature sequence; Group and label several groups of fusion feature sequences, denoted as which is a natural number; Take groups of several groups of fusion feature sequences as sample data, and is a natural number less than , and use the sample data to obtain the sample data mean, denoted as the sample set; Take the remaining several groups of fusion feature sequences and the standard fusion feature sequence as the test set; According to the sample set and the test set, form a training sample set; Based on the convolutional neural network, construct a standard sound classification model; And input the training sample set into the standard sound classification model, train the standard sound classification model, obtain the trained standard sound classification model, and denote the trained standard sound classification model as the sound classification target model.
[0014] According to one preferred embodiment of the present invention, the process of generating a dynamic sound index based on the sound classification target model and then obtaining the sound classification result includes: Generate a dynamic sound index under the current environmental factor conditions according to the sound classification target model , and the dynamic sound index is: ; wherein, represents the actual environmental temperature, represents the standard environmental temperature, represents the proportional coefficient of the temperature influence degree, represents the maximum amplitude value of the sensor vibration sound data, Represents the texture feature similarity between the local area of the i-th pixel and the local area of the j-th pixel. Represents the current standard environmental factors. Represents the actual environmental factors. Preset standard dynamic sound index ; If the dynamic sound index , then the sound under the current environmental factor conditions conforms to the classification standard. If the dynamic sound index , then the sound under the current environmental factor conditions does not conform to the classification standard.
[0015] The second aspect of the present invention further provides a system for constructing a sound classification target model. The system executes the above-mentioned method for constructing a sound classification target model, including: a sound data acquisition module, a sound data processing module, a sound data feature processing module, a sound data storage module, and a classification target model construction module.
[0016] The sound data acquisition module is used to acquire original sound data and auxiliary modal data. The sound data processing module is used to process the original sound data and auxiliary modal data to obtain multi-resolution sound image data and an auxiliary modal spatio-temporal matrix. The sound data feature processing module is used to perform local feature processing on the multi-resolution sound image data and the auxiliary modal spatio-temporal matrix to obtain local sound image features and auxiliary modal features, and construct a hierarchical feature sequence space to obtain a fused feature sequence. The sound data storage module is used to store the fused feature sequence. The classification target model construction module is used to construct a sound classification target model according to the fused feature sequence, generate a dynamic sound index according to the sound classification target model, and further obtain a sound classification result.
[0017] The present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement the above-mentioned method for constructing a sound classification target model.
[0018] Compared with the prior art, the beneficial effects of the present invention are: acquiring original sound data and auxiliary modal data, performing sound data conversion processing on the original sound data to obtain corresponding multi-resolution sound image data; processing the auxiliary modal data to obtain a corresponding auxiliary modal spatio-temporal matrix; dynamically adapting to the importance changes of each modality in a complex environment, and improving the sound recognition ability.
[0019] Perform local feature processing on the multi-resolution sound image data and the auxiliary modality spatio-temporal matrix to obtain local sound image features and auxiliary modality features; construct a hierarchical feature sequence space based on the local sound image features and the auxiliary modality features to obtain a corresponding fused feature sequence; construct a sound classification target model based on the fused feature sequence, and generate a dynamic sound index according to the sound classification target model, thereby obtaining a sound classification result. After training under laboratory conditions and deploying it to a real environment (such as the ocean, industrial site), it is less susceptible to factors such as noise, reverberation, and sensor differences, and improves the accuracy of sound classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a schematic diagram of a method for constructing a sound classification target model.
[0022] Figure 2 It is a module diagram of a system for constructing a sound classification target model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations. The basic principles defined in the following description can be applied to other implementation schemes, variation schemes, improvement schemes, equivalent schemes, and other technical schemes that do not deviate from the spirit and scope of the present invention.
[0024] It can be understood that the term "one" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, and in other embodiments, the number of the element can be multiple. The term "one" cannot be understood as a limitation on the number.
[0025] As Figure 1 shown, a method for constructing a sound classification target model includes the following steps: Obtain the original sound data and the auxiliary modality data, perform sound data conversion processing on the original sound data to obtain corresponding multi-resolution sound image data; process the auxiliary modality data to obtain a corresponding auxiliary modality spatio-temporal matrix; Perform local feature processing on the multi-resolution sound image data and the auxiliary modality spatio-temporal matrix to obtain local sound image features and auxiliary modality features; Construct a hierarchical feature sequence space based on the local sound image features and auxiliary modality features to obtain the corresponding fused feature sequence; Construct a sound classification target model according to the fused feature sequence, and generate a dynamic sound index according to the sound classification target model, thereby obtaining a sound classification result.
[0026] It should be further noted that in the specific implementation process, the specific process of obtaining the original sound data and auxiliary modality data includes: Preset a data acquisition device, which is composed of several data acquisition units, and set an acquisition period. The data acquisition units include an original sound data acquisition unit, a device vibration data acquisition unit, and an ambient temperature data acquisition unit; The original sound data acquisition unit is used to acquire original sound data; the device vibration data acquisition unit is used to acquire the sensor vibration sound data when the data acquisition unit is working; the ambient temperature data acquisition unit is used to acquire the ambient temperature during the current acquisition period. It should be further noted that during the process of acquiring sound data by the data acquisition device, the corresponding data acquisition unit will generate subtle vibration sound data; the sensor vibration sound data and the ambient temperature are recorded as auxiliary modality data.
[0027] It should be further noted that in the specific implementation process, the specific process of performing sound data conversion processing on the original sound data includes: Obtain the original sound data; Perform preprocessing on the original sound data. The specific process includes: Select a sampling frequency of 44100 Hz and a quantization bit number of 24 bits, and perform sampling and quantization on the original sound data. It should be further noted that select appropriate sampling frequency and quantization bit number according to actual requirements; Remove the noise interference in the original sound data after sampling and quantization based on a filter, and retain the frequency components related to sound classification; Normalize the amplitude of the original sound data after removing the noise interference to [-1, 1] to avoid affecting subsequent processing due to excessive signal amplitude differences; Based on the Hamming window function, divide the preprocessed original sound data into multiple short time segments, and perform Fourier transform on each segment to obtain a time-frequency representation , the time-frequency representation is: ; where is the original sound data, is the Hamming window function, is the frame shift, is the window length, m is the frame index, and k is the frequency index; Adjust the window length and frame shift of the Hamming window function to generate Mel spectrograms with different resolutions, including Mel spectrograms with resolutions of 256×256 and 128×128; It should be further noted that Mel spectrograms with different resolutions can capture features of different scales of audio signals. High-resolution spectrograms can retain more detailed information, while low-resolution spectrograms can highlight the overall feature trends; Based on the Griffin-Lim algorithm, estimate the phase information from the time-frequency representation Initially, randomly initialize the phase, and then synthesize a signal from the time-frequency representation and the current initialized phase, and then perform STFT calculation on the synthesized signal to obtain a new time-frequency representation and phase, continuously iterate and update the phase until the convergence condition is met; Obtain a high-resolution spectrogram with higher accuracy; Optimize the high-resolution spectrogram based on the Wasserstein GAN technology. By minimizing the Wasserstein distance, improve the quality of the high-resolution spectrogram, and denote it as multi-resolution sound image data.
[0028] It should be further noted that in the specific implementation process, the specific process of processing the auxiliary modal data includes: Preset several feature extraction times; Based on the filtering processing technology, remove the noise interference of the sensor vibration sound data and retain the useful signal components. Based on the wavelet denoising technology, improve the quality of the filtered sensor vibration sound data; Extract features from the sensor vibration sound data after filtering and denoising, specifically including: 1. Amplitude feature extraction, obtain the maximum value, mean value, and root mean square value of the sensor vibration sound data within the feature extraction time; 2. Frequency feature extraction, based on the Fourier transform, convert the sensor vibration sound data within the feature extraction time to the frequency domain, and calculate the main frequency and spectral energy distribution of the sensor vibration sound data; 3. Energy feature extraction, perform wavelet packet decomposition on the sensor vibration sound data within the feature extraction time, and calculate the energy of each wavelet packet node; Synthesize a matrix according to the amplitude feature extraction, frequency feature extraction, and energy feature extraction at different feature extraction times to form a spatio-temporal matrix of the sensor vibration sound data; Preset feature extractions are performed on the sensor vibration sound data at time points, and each time point extracts Generate a spatio-temporal matrix of sensor vibration sound data for each feature , where the spatio-temporal matrix is: ; Arrange the environmental temperature in the order of the feature extraction time to form a one-dimensional vector , ensuring the correct time order of the data for matching with the sensor vibration sound data; Generate a corresponding auxiliary modal spatio-temporal matrix and one-dimensional vector , where , , are weight coefficients.
[0029] It should be further noted that in the specific implementation process, the specific process of performing local feature processing on the multi-resolution sound image data and the auxiliary modal spatio-temporal matrix includes:
[0030] The process of generating the local sound image feature set of the multi-resolution sound image data based on the gray-level co-occurrence matrix technology includes: Calculate the gray-level co-occurrence matrix around each pixel point of the multi-resolution sound image data by means of a sliding window. The specific details are as follows. Quantify the gray-level of the multi-resolution sound image data into discrete gray-levels, divide the gray-value range into several levels. For example, an 8-bit image can be divided into 16, 32, 64 levels, and define the parameters required for the gray-level co-occurrence matrix, including distance and direction, which are determined according to the actual application requirements. The specific steps will not be elaborated. Subsequently, based on the calculated gray-level co-occurrence matrix, extract a series of local sound image features and construct a local sound image feature set. The local sound image features in the local sound image feature set include but are not limited to: Contrast: A statistical feature describing the contrast of pixels with different gray-levels in the image, a feature measuring the roughness of the image texture, reflecting the clarity of the image and the depth of the texture grooves. The deeper the texture grooves, the greater the contrast and the clearer the visual effect. On the contrary, when the contrast is small, the grooves are shallow and the effect is blurred; Correlation: Describing the degree of correlation of pixels with different gray-levels in the image, it measures the similarity degree of the elements of the spatial gray-level co-occurrence matrix in the row or column direction. Therefore, the size of the correlation value reflects the local gray-level correlation in the image. When the matrix element values are uniformly equal, the correlation value is large. On the contrary, if the matrix pixel values vary greatly, the correlation value is small. If there are horizontal texture in the image, the correlation of the horizontal matrix is greater than that of the other matrices; Energy: Describes the uniformity of the gray - level distribution of pixels in an image, measures the randomness contained in the image, and represents the complexity of the image. When all values in the co - occurrence matrix are equal or the pixel values show the maximum randomness, the entropy is the largest; Homogeneity: Describes the similarity degree of gray - level of adjacent pixels in an image; reflects the homogeneity of the image texture, measures the amount of local change in the image texture. A large value indicates that there is little change between different regions of the image texture and it is very uniform locally; Entropy: Describes the degree of uncertainty of the image texture, measures the randomness contained in the image, and represents the complexity of the image. When all values in the co - occurrence matrix are equal or the pixel values show the maximum randomness, the entropy is the largest; Inverse variance: Reflects the clarity and regularity of the texture. The texture is clear and has strong regularity; It should be further noted that in the specific implementation process, the calculation formula for obtaining the texture feature similarity between the local area of each pixel point and the local area of other adjacent pixel points is: ; Among them, represents the texture feature similarity between the local area of the i - th pixel point and the local area of the j - th pixel point, represents the contrast between the local area of the i - th pixel point and the local area of the j - th pixel point, represents the energy of the local area of the i - th pixel point, represents the energy of the local area of the j - th pixel point, represents the entropy of the local area of the i - th pixel point, represents the entropy of the local area of the j - th pixel point, represents the inverse variance of the local area of the i - th pixel point, represents the inverse variance of the local area of the j - th pixel point, , , , and represent weight factors.
[0031] Perform a linear mapping on the auxiliary modal spatio - temporal matrix. The specific process is as follows: Obtain the feature dimension of the auxiliary modal spatio - temporal matrix ; It should be further noted that the values of the feature dimension are 128, 256, 512, etc., and are selected according to actual needs; Preset the target feature dimension ; Map the feature dimension of the auxiliary modal spatio - temporal matrix to the target dimension , and the calculation process of the linear mapping is as follows: ; where, is the weight matrix, with a shape of , b is the bias vector, with a shape of , is the output matrix after mapping, with a shape of , is the auxiliary modality spatio-temporal matrix 's batch size, is the auxiliary modality spatio-temporal matrix 's sequence length; Based on the Transformer encoder, learn the output matrix to output the auxiliary modality features.
[0032] It should be further noted that in the specific implementation process, the specific process of constructing the hierarchical feature sequence space includes: Obtain the local sound image feature set and the auxiliary modality features; Fuse the local sound image feature set and the auxiliary modality features to obtain a fused feature set, denoted as , and the fused feature set is: ; where, is the local sound image feature set, is the auxiliary modality feature, , are the weight coefficients; Based on the small neural network technology, dynamically allocate the weight coefficients of the local sound image feature set and the auxiliary modality features. It should be further noted that the specific allocation of the weight coefficients can be modified according to the actual specific situation; According to the fused feature set and the dynamically allocated weight coefficients, generate the hierarchical feature sequence space and then the corresponding fused feature sequence.
[0033] It should be further noted that in the specific implementation process, the specific process of constructing the sound classification target model includes: Obtain several groups of fused feature sequences; Preset the standard fused feature sequence; it should be further noted that the standard fused feature sequence includes the standard fused feature sequences of environmental factors such as noise and reverberation; Group and label several groups of fused feature sequences, denoted as is a natural number; Take groups of several groups of fused feature sequences as sample data, and is less than natural numbers, and using the sample data, obtain the mean of the sample data, denoted as the sample set; Use the remaining several groups of fusion feature sequences and the standard fusion feature sequence as the test set; According to the sample set and the test set, form a training sample set; Based on the convolutional neural network, construct a standard sound classification model; And input the training sample set into the standard sound classification model, train the standard sound classification model, obtain the trained standard sound classification model, and denote the trained standard sound classification model as the sound classification target model.
[0034] It should be further noted that in the specific implementation process, according to the sound classification target model, generating a dynamic sound index and then obtaining the specific process of the sound classification result includes: According to the sound classification target model, generate a dynamic sound index under the current environmental factor conditions , the dynamic sound index is: ; where represents the actual environmental temperature, represents the standard environmental temperature, represents the proportionality coefficient of the temperature influence degree, represents the maximum amplitude value of the sensor vibration sound data, represents the texture feature similarity between the local area of the i-th pixel point and the local area of the j-th pixel point, represents the current standard environmental factor, represents the actual environmental factor; Preset the standard dynamic sound index ; If the dynamic sound index , then the sound under the current environmental factor conditions meets the classification standard; If the dynamic sound index , then the sound under the current environmental factor conditions does not meet the classification standard.
[0035] As Figure 2 shown, a sound classification target model construction system includes: a sound data acquisition module, a sound data processing module, a sound data feature processing module, a sound data storage module, and a classification target model construction module; The sound data acquisition module is used to acquire the original sound data and the auxiliary modality data; The sound data processing module is used to process the original sound data and the auxiliary modality data to obtain the multi-resolution sound image data and the auxiliary modality spatio-temporal matrix; The voice data feature processing module is used to perform local feature processing on the multi-resolution voice image data and the auxiliary modality spatio-temporal matrix, obtain local voice image features and auxiliary modality features, construct a hierarchical feature sequence space, and obtain a fused feature sequence; The voice data storage module is used to store the fused feature sequence; The classification target model construction module is used to construct a voice classification target model according to the fused feature sequence, generate a dynamic voice index according to the voice classification target model, and then obtain a voice classification result.
[0036] In the embodiments disclosed by the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. The embodiments disclosed by the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed. It should be noted that the above computer-readable medium in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wire segments, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless segments, wire segments, optical cables, RF, etc., or any suitable combination of the above.
[0037] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0038] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and illustrated in the embodiments. Without departing from the said principles, the embodiments of the present invention may have any variations or modifications.
Claims
1. A method for constructing a sound classification target model, characterized in that: include: Acquire original sound data and auxiliary modal data, perform sound data conversion processing on the original sound data, and obtain corresponding multi-resolution sound image data; Processing the auxiliary modality data to obtain a corresponding auxiliary modality space-time matrix; Performing local feature processing on the multi-resolution sound image data and the auxiliary modality spatiotemporal matrix to obtain local sound image features and auxiliary modality features; According to the local sound image features and the auxiliary modality features, a hierarchical feature sequence space is constructed to obtain a corresponding fusion feature sequence; According to the fusion feature sequence, a sound classification target model is constructed, and according to the sound classification target model, a dynamic sound index is generated to obtain a sound classification result.
2. A method for constructing a sound classification target model according to claim 1, characterized in that: The process of obtaining the original sound data and auxiliary modality data includes: A data acquisition device is preset, the data acquisition device is composed of a plurality of data acquisition units, and a collection cycle is set, the data acquisition units include an original sound data acquisition unit, an equipment vibration data acquisition unit and an ambient temperature data acquisition unit; The original sound data collection unit is used to collect original sound data; The equipment vibration data acquisition unit is used to collect sensor vibration sound data when the data acquisition unit is working; The ambient temperature data acquisition unit is used to acquire the ambient temperature in the current acquisition cycle; The sensor vibration sound data and the ambient temperature are recorded as auxiliary modal data.
3. A method for constructing a sound classification target model according to claim 2, characterized in that: The process of converting the original sound data into sound data includes: Get the original sound data; The original sound data is preprocessed, and the preprocessed original sound data is divided into multiple short time periods based on the Hamming window function, and each time period is subjected to Fourier transform to obtain a time-frequency representation. , the time-frequency representation for: ;in, is the original sound data, is the Hamming window function, is the frame shift, is the window length, m is the frame index, and k is the frequency index; Based on the Griffin-Lim algorithm, obtain a more accurate high-resolution spectrum diagram; Based on the Wasserstein GAN technology, the high-resolution spectrogram is optimized to improve the quality of the high-resolution spectrogram and recorded as multi-resolution sound image data.
4. A method for constructing a sound classification target model according to claim 3, characterized in that: The process of processing auxiliary modality data includes: Preset several feature extraction times; Based on filtering processing technology, the noise interference of sensor vibration sound data is removed, the useful signal components are retained, and based on wavelet noise reduction technology, the quality of filtered sensor vibration sound data is improved; Perform amplitude feature extraction, frequency feature extraction and energy feature extraction on the sensor vibration sound data after filtering and noise reduction, and synthesize matrices according to the amplitude feature extraction, frequency feature extraction and energy feature extraction at different feature extraction times to form a spatiotemporal matrix of the sensor vibration sound data; Presets The sensor vibration sound data was feature extracted at each time point. Features are generated to generate a spatiotemporal matrix of sensor vibration and sound data; Arrange the ambient temperatures in the order of the feature extraction time to form a one-dimensional vector; A corresponding auxiliary modal space-time matrix is generated according to the space-time matrix and the one-dimensional vector.
5. A method for constructing a sound classification target model according to claim 4, characterized in that: The process of local feature processing of multi-resolution sound image data and auxiliary modal space-time matrix includes: Based on the gray level co-occurrence matrix technology, a local sound image feature set of multi-resolution sound image data is generated, and the texture feature similarity between the local area of each pixel point and the local areas of other adjacent pixel points is obtained according to the local sound image feature set; The auxiliary modal space-time matrix is linearly mapped, and the process is: Get the characteristic dimension of the auxiliary modal space-time matrix ; Preset target feature dimensions ; The characteristic dimension of the auxiliary modal space-time matrix Mapping to target dimension , the calculation process of linear mapping is as follows: ;in, is the space-time matrix, is the weight matrix, with the shape , b is the bias vector, the shape is , is the output matrix after mapping, the shape is , is the auxiliary modal space-time matrix The batch size is is the auxiliary modal space-time matrix The length of the sequence; Based on the Transformer encoder, and according to the output matrix , output auxiliary modal features.
6. A method for constructing a sound classification target model according to claim 5, characterized in that: The process of constructing a hierarchical feature sequence space includes: Obtaining a local sound image feature set and auxiliary modal features; The local sound image feature set and the auxiliary modal features are fused to obtain a fused feature set, which is recorded as , the fusion feature set for: ;in, is the local sound image feature set, is the auxiliary modal feature, , is the weight coefficient; Based on small neural network technology, the weight coefficients of local sound image feature sets and auxiliary modal features are dynamically allocated; According to the fusion feature set and dynamically assigned weight coefficients to generate a hierarchical feature sequence space and the corresponding fused feature sequence.
7. A method for constructing a sound classification target model according to claim 6, characterized in that: The process of building a sound classification target model includes: Obtaining several groups of fusion feature sequences; Preset standard fusion feature sequence; Group and label several groups of fusion feature sequences, denoted as is a natural number; Will Several groups of fusion feature sequences are used as sample data, and is less than A natural number, and using the sample data, obtain the mean of the sample data, recorded as a sample set; The remaining groups of fused feature sequences and standard fused feature sequences are used as test sets; According to the sample set and the test set, a training sample set is formed; Based on convolutional network neural network, a standard sound classification model is constructed; The training sample set is input into the standard sound classification model, the standard sound classification model is trained, the trained standard sound classification model is obtained, and the trained standard sound classification model is recorded as the sound classification target model.
8. A method for constructing a sound classification target model according to claim 7, characterized in that: The process of generating a dynamic sound index according to the sound classification target model and then obtaining the sound classification result includes: Generate a dynamic sound index under current environmental factors according to the sound classification target model , the dynamic sound index for: ;in, Indicates the actual ambient temperature. Indicates the standard ambient temperature, The proportionality factor indicating the degree of temperature influence, Indicates the maximum amplitude value of the sensor vibration sound data. represents the texture feature similarity between the local area of the i-th pixel and the local area of the j-th pixel, Indicates the current standard environmental factors, Indicates actual environmental factors; Preset standard dynamic sound index ; If the dynamic sound index , then the sound under the current environmental factors meets the classification criteria; If the dynamic sound index , then the sound under the current environmental factors does not meet the classification criteria.
9. A sound classification target model construction system, the system executing a sound classification target model construction method according to any one of claims 1 to 8, comprising: Sound data acquisition module, sound data processing module, sound data feature processing module, sound data storage module and classification target model building module; The sound data acquisition module is used to collect original sound data and auxiliary modal data; The sound data processing module is used to process the original sound data and the auxiliary modal data to obtain multi-resolution sound image data and the auxiliary modal space-time matrix; The sound data feature processing module is used to perform local feature processing on the multi-resolution sound image data and the auxiliary modality spatiotemporal matrix to obtain local sound image features and auxiliary modality features, and to construct a hierarchical feature sequence space to obtain a fused feature sequence; The sound data storage module is used to store the fusion feature sequence; The classification target model construction module is used to construct a sound classification target model according to the fusion feature sequence, generate a dynamic sound index according to the sound classification target model, and then obtain a sound classification result.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program can be executed by a processor to implement a method for constructing a sound classification target model as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Voice-based video action classification method and related equipment
CN114529846A
Voiceprint recognition method of DenseNet-LSTM-ED based on global attention mechanism
CN116863939A
Substation transformer fault infrared image and voiceprint recognition method and system
CN118051873A
Structural damage identification method and system in combination with multi-modal information and artificial intelligence
CN118552795A
Cited By
Marine mammal sound classification method based on IVGG-ASNet
CN120260588A
Drawing machine fault diagnosis method and system based on multi-source data fusion
CN120763836A
Underwater sound target identification method based on Wasserstein vector spectrum
CN121034349A