Machine voiceprint feature modeling method, system, device, medium and program product

By performing time segmentation and frequency importance evaluation on the two-dimensional time-frequency spectrogram, the problem of insufficient machine voiceprint feature extraction capability in existing technologies is solved, a more robust and generalized voiceprint feature representation is achieved, and the adaptability and recognition accuracy of the fault diagnosis model are improved.

CN120766686AActive Publication Date: 2025-10-10NORTHEASTERN UNIV CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511144694.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-10
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively capture the key time-frequency information of machine voiceprints in complex and changing industrial environments, resulting in insufficient accuracy and generalization of fault diagnosis models.

Method used

By performing time-block processing on the two-dimensional time-frequency spectrogram and introducing a frequency importance evaluation mechanism and a sorting pooling strategy, the voiceprint characteristics of key time and frequency frames are extracted and a machine voiceprint feature matrix is ​​constructed.

Benefits of technology

It significantly improves the accuracy and robustness of industrial machine voiceprint modeling, adapts to different machine types and working conditions, and improves the recognition accuracy and generalization ability of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766686A_ABST
    Figure CN120766686A_ABST
Patent Text Reader

Abstract

The invention provides a machine voiceprint feature modeling method, system and device, a medium and a program product, and relates to the technical field of voiceprint analysis. The invention provides a machine voiceprint feature modeling method aiming at the problems of audio feature extraction and representation modeling methods, which comprises the following steps: dividing a time-frequency spectrogram into a plurality of sub-blocks in a time dimension, sequentially weighting each sub-block, introducing a frequency importance evaluation mechanism and a sorting pooling strategy, and carrying out feature extraction on the sub-blocks; performing significance weighting on the frequency dimension features in each sub-block, and sequentially splicing weighted feature vectors of each sub-block to obtain a machine voiceprint feature matrix; according to the method, time partitioning processing is performed on the two-dimensional time-frequency spectrogram, a frequency importance evaluation mechanism and a sorting pooling strategy are introduced, the voiceprint characteristics of key time and frequency frames are fully captured and learned, and the performance and the adaptive capacity of industrial machine voiceprint modeling in an actual fault diagnosis task are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of voiceprint analysis technology, and specifically relates to a machine voiceprint feature modeling method, system, device, medium and program product. Background Art

[0002] With the increasing level of industrial automation, the complexity of industrial machines and the changing operating environments are placing higher demands on real-time monitoring and intelligent diagnosis of equipment operating status. Traditional fault diagnosis methods often rely on manual experience, are inefficient, and are susceptible to subjective factors. In recent years, with the rapid development of signal processing and artificial intelligence technologies, industrial machine voiceprint analysis has gradually become a key means of addressing this problem. The so-called "machine voiceprint" refers to the unique sound pattern emitted by a machine under specific structure, working conditions, and operating conditions. By extracting the voiceprint characteristics of such patterns, it can be used to identify whether the equipment is in a normal state or experiencing potential anomalies.

[0003] However, due to the diverse nature of industrial machines in real-world production, their complex operating states, and the significant differences in the sound spectral structures of different devices, coupled with challenges such as environmental noise interference and scarce fault data, it is difficult to effectively characterize machine voiceprints using only traditional audio time- and frequency-domain features (such as RMS value, short-time energy, Mel-frequency cepstral coefficients, and power spectral density) in this domain-shifted scenario. To address these issues, current approaches often convert one-dimensional sound signals into two-dimensional time-frequency spectrograms and extract their statistical features to model and classify voiceprint features. Alternatively, neural networks are used to directly extract more complex and abstract feature representations of the original sound signal in an end-to-end manner. These methods can improve voiceprint analysis capabilities to a certain extent. However, these methods generally use global pooling or statistical averaging in the feature aggregation stage, failing to dynamically model local variations in the time-frequency spectrum and highlight key time periods and important frequency regions, limiting their applicability and generalization in complex and variable real-world operating conditions.

[0004] Therefore, designing an effective machine voiceprint feature representation and modeling method is of great significance for improving the accuracy and robustness of industrial voiceprints in tasks such as anomaly detection, fault diagnosis and identification. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a machine voiceprint feature modeling method, system, equipment, medium and program product to solve the problem in actual industry that there is a lack of feature representation technology that can efficiently and fully capture the key time-frequency information of machine voiceprints, thereby improving the accuracy and generalization of existing fault diagnosis models in identifying the operating status of industrial machines and detecting anomalies in complex environments.

[0006] In a first aspect, the present invention provides a method for modeling machine voiceprint features, comprising the following steps:

[0007] Step 1: Collect the sound signals of industrial equipment operation, pre-process the sound signals of industrial equipment operation, and generate a two-dimensional time-frequency spectrogram;

[0008] The preprocessing of industrial equipment operation sound signals includes: pre-emphasis, framing, windowing, denoising and short-time Fourier transform operations on the industrial equipment operation sound signals, and using Mel transform to convert the one-dimensional audio signal into a two-dimensional time-frequency spectrogram , where F is the frequency dimension and T is the time dimension;

[0009] Step 2: Divide the two-dimensional time-frequency spectrogram into several non-overlapping sub-blocks;

[0010] Divide the two-dimensional time-frequency spectrogram X into N non-overlapping sub-blocks of length W along the time dimension T. , , where the time length W=T / N;

[0011] Step 3: Sort the time frames in each sub-block of the two-dimensional time-frequency spectrogram by their importance based on the time dimension, and then convert each sub-block of the two-dimensional time-frequency spectrogram into a one-dimensional vector. Perform weighted pooling on the sorted sub-blocks to obtain a one-dimensional feature vector for each sub-block.

[0012] Step 3.1: For each sub-block of the two-dimensional time-frequency spectrogram, sort the energy of the frequency frame of each sub-block according to the time dimension. After sorting, the two-dimensional time-frequency spectrogram is ;

[0013] Step 3.2: Use learnable pooling vectors For the sorted sub-blocks Perform weighted pooling to obtain a one-dimensional feature vector that characterizes the voiceprint characteristics of each sub-block in a local time period. , as shown in the following formula:

[0014]

[0015]

[0016] Where T represents the matrix transpose calculation, r∈[0,1] is a learnable parameter, is the sum of the learnable parameters r;

[0017] Step 4: Calculate the frequency attention weight of the two-dimensional time-frequency spectrogram;

[0018] Step 4.1: Calculate the frequency importance value of the two-dimensional time-frequency spectrogram;

[0019] Calculate the global average value of each frequency segment f∈F in the time dimension T of the two-dimensional time-frequency spectrogram X as the frequency importance value to measure the overall energy level of the frequency segment , as shown in the following formula:

[0020]

[0021] Frequency importance value The larger the value, the more important the frequency segment f is;

[0022] Step 4.2: Calculate the frequency attention weight of the two-dimensional time-frequency spectrogram based on the frequency importance value of the two-dimensional time-frequency spectrogram;

[0023] Use a multi-layer perceptron network and calculate the weight vector of the frequency importance value of each frequency segment through the Softmax normalization layer , as shown in the following formula:

[0024]

[0025] Among them, K and b are learnable parameters, and ;

[0026] Step 5: Calculate the weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram based on the frequency attention weight of the two-dimensional time-frequency spectrogram and the one-dimensional feature vector of each sub-block;

[0027] The weight vector A of the frequency importance value of each frequency segment f Sequentially with the one-dimensional feature vector of each sub-block Multiply element by element and weight the two-dimensional time-frequency spectrogram according to the frequency importance to obtain the weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram , as shown in the following formula:

[0028]

[0029] Step 6: Concatenate the weighted feature vectors of each sub-block in sequence to obtain the machine voiceprint feature matrix;

[0030] The feature vector of each sub-block Splice them in sequence to get the machine voiceprint feature matrix , as shown in the following formula:

[0031] .

[0032] In a second aspect, the present invention further provides a machine voiceprint feature modeling system, comprising: an industrial sound signal data acquisition module, a sound preprocessing and denoising module, and a machine voiceprint characterization and modeling module;

[0033] The industrial sound signal data acquisition module is used to collect the sound signals of industrial equipment operation, store the collected industrial equipment operation sound signals in the form of audio signals, and transmit them to the sound preprocessing and denoising module;

[0034] The sound preprocessing and denoising module is used to preprocess the collected sound signals of industrial equipment operation. It uses Mel transform to convert the one-dimensional audio signal into a two-dimensional time-frequency spectrogram, and then transmits the two-dimensional time-frequency spectrogram to the machine voiceprint representation and modeling module.

[0035] The machine voiceprint representation and modeling module is used to process the two-dimensional time-frequency spectrogram, divide the two-dimensional time-frequency spectrogram into several sub-blocks in the time dimension, and extract the one-dimensional feature vector of each sub-block; at the same time, the weighted feature vector of each sub-block is calculated based on the frequency attention weight; finally, the weighted feature vectors of each sub-block are spliced ​​in sequence to obtain the machine voiceprint feature matrix.

[0036] In a third aspect, the present application proposes an electronic device comprising: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the machine voiceprint feature modeling method.

[0037] In a fourth aspect, the present application proposes a computer-readable storage medium storing executable instructions, which, when executed, enable a processor to execute the machine voiceprint feature modeling method.

[0038] In a fifth aspect, the present application proposes a computer program product, comprising a computer program or instructions, which implement the machine voiceprint feature modeling method when executed by a processor.

[0039] The beneficial effects of adopting the above technical solution are as follows: the machine voiceprint feature modeling method provided by the present invention addresses the problems existing in existing industrial voiceprint feature extraction methods, such as weak key information extraction capability, insufficient time-frequency resolution, and poor model generalization capability. It enhances the modeling capability of local voiceprint features, improves the recognition and diagnostic effectiveness of the frequency dimension, and improves the adaptability and robustness of the diagnostic model. It has strong adaptability and broad application prospects and practical value. Specifically:

[0040] By dividing the original time-frequency spectrogram into multiple time blocks and performing weighted pooling and feature aggregation within each local block, the model's ability to identify local key features is significantly improved, avoiding the problem of feature representation ambiguity and neglect of dynamic time information caused by basic methods;

[0041] The introduction of a frequency importance assessment mechanism enables the voiceprint characterization modeling process to focus more on frequency regions that are strongly related to the device's operating status, improving recognition accuracy under abnormal conditions.

[0042] Through a unified structured strategy of segmentation, adaptive weighting, and aggregation, it is possible to efficiently capture the voiceprint characteristics of different machine types and learn more robust and higher-level audio representations, thereby further improving the recognition accuracy and generalization ability of subsequent diagnostic models in non-ideal acoustic environments such as multi-state and multi-device types.

[0043] The present invention has a simple structure and is explainable. It can be flexibly applied to the task of operating status voiceprint modeling of various types of industrial equipment, including power equipment, machine tools, fans, pumps and valves. It has good engineering adaptability and deployment feasibility, and can provide reliable voiceprint modeling support for the field of industrial equipment fault detection and diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A flowchart of the machine voiceprint block-based time-frequency domain weighted pooling representation modeling process provided by Example 1 of the present invention;

[0045] Figure 2 Schematic diagram of the machine voiceprint feature modeling system structure provided by Example 2 of the present invention;

[0046] Figure 3 Example 3 of the present invention provides a modeling result of the voiceprint features of the sound signal of the transformer under abnormal operation state, wherein (a) is the original spectrogram of the sound signal of the transformer under abnormal operation state, (b) is the sorted two-dimensional time-frequency spectrum of the sound signal of the transformer under abnormal operation state, (c) is the pooling vector, and (d) is the machine voiceprint feature matrix of the sound signal of the transformer under abnormal operation state. DETAILED DESCRIPTION

[0047] The specific implementation of the present application is further described in detail below with reference to the accompanying drawings and examples.

[0048] Example 1

[0049] In the field of industrial machine voiceprint analysis, several methods exist for extracting and modeling sound features. The most commonly used technique is to convert the original audio signal into a two-dimensional time-frequency spectrum through the Short-time Fourier Transform (STFT). The entire time-frequency spectrum is then statistically averaged or variance-modeled to construct a feature vector. While these methods are simple to implement and computationally efficient, their feature representation process fails to fully highlight the differential characteristics of key time segments and important frequency regions. In recent years, deep learning-based methods have been increasingly used for audio feature representation. For example, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory (LSTMs) have been used to extract deep audio features. However, these neural network structures are complex, the extracted features lack interpretability, and obtaining an ideal model requires extensive computing resources and training data, both of which are unavailable in real industrial environments. Another approach involves performing a global weighted average pooling operation on the converted two-dimensional time-frequency spectrum along the time dimension, generating a one-dimensional pooling vector to represent the machine's voiceprint. However, the acoustic signals generated by operating machinery exhibit strong non-stationarity and diversity. Methods that rely on a unified global modeling of the entire signal cannot capture information between time segments, resulting in the neglect of short-term dynamic features and an inability to adapt to the complex and changing working conditions of real applications.

[0050] These audio feature extraction and representation modeling methods mentioned above still have some problems, such as ignoring time domain dynamic anomaly information, failing to perform differentiated modeling of key frequency areas, and lack of interpretability in the representation learning process.

[0051] This embodiment proposes a machine voiceprint feature modeling method. By performing time-block processing on the two-dimensional time-frequency spectrogram and introducing a frequency importance evaluation mechanism and a sorting pooling strategy, it fully captures and learns the voiceprint characteristics of key time and frequency frames, effectively improving the performance and adaptability of industrial machine voiceprint modeling in actual fault diagnosis tasks.

[0052] The machine voiceprint feature modeling method of this embodiment is as follows: Figure 1 As shown, the following steps are included:

[0053] Step 1: Collect the sound signals of industrial equipment operation, pre-process the sound signals of industrial equipment operation, and generate a two-dimensional time-frequency spectrogram;

[0054] The preprocessing of the industrial equipment running sound signal comprises: pre-emphasis, frame division, windowing, denoising and short-time Fourier transform operation on the industrial equipment running sound signal, and converting a one-dimensional audio signal into a two-dimensional time-frequency spectrogram by using a Mel transform , wherein F is a frequency dimension, and T is a time dimension;

[0055] Step 2: dividing the two-dimensional time-frequency spectrogram into a plurality of non-overlapping sub-blocks;

[0056] The two-dimensional time-frequency spectrogram X is evenly divided into N non-overlapping sub-blocks along the time dimension T , , wherein the time length W = T / N;

[0057] Step 3: sequentially performing time-dimension-based importance sorting on time frames in each sub-block of the two-dimensional time-frequency spectrogram, thereby converting each sub-block of the two-dimensional time-frequency spectrogram into a one-dimensional vector, performing weighted pooling on the sorted sub-blocks, and obtaining a one-dimensional feature vector of each sub-block;

[0058] Step 3.1: for each sub-block of the two-dimensional time-frequency spectrogram, sorting the energy of the frequency frames of each sub-block according to the time dimension, and the sorted two-dimensional time-frequency spectrogram is ;

[0059] Step 3.2: using a learnable pooling vector to perform weighted pooling on the sorted sub-blocks , thereby obtaining a one-dimensional feature vector for representing the voiceprint characteristics of each sub-block in a local time period , as shown in the following formula:

[0060]

[0061]

[0062] , wherein T represents matrix transposition calculation, r [0, 1] is a learnable parameter, is the sum of the learnable parameter r;

[0063] When the learnable parameter r = 0, it represents maximum pooling calculation, at this time, the stationary characteristics of the running state are ignored; when the learnable parameter r = 1, it represents average pooling calculation, at this time, the transient characteristics of the running state are ignored; according to the running characteristics of different industrial equipment, the learnable parameter r is adaptively set to different values, compared with the global sorting pooling method, the time information of the two-dimensional time-frequency spectrogram X can be better preserved, and the machine voiceprint characteristics can be more robustly represented and modeled, and the one-dimensional feature vector of the sorted sub-block is lower in dimension than the two-dimensional time-frequency spectrogram X ;

[0064] Step 4: Calculate the frequency attention weight of the two-dimensional time-frequency spectrogram;

[0065] The weighted pooling operation in step 3 increases reliance on the time dimension but fails to consider the importance of different frequency bands for different machine types and operating conditions. These different frequency bands are crucial for anomaly monitoring and fault diagnosis of industrial equipment. By extracting frequency importance scores and assigning different weights to different frequency bands, the classifier focuses more on frequency regions that are critical and important for fault detection.

[0066] Step 4.1: Calculate the frequency importance value of the two-dimensional time-frequency spectrogram;

[0067] Calculate the global average value of each frequency segment f∈F in the time dimension T of the two-dimensional time-frequency spectrogram X as the frequency importance value to measure the overall energy level of the frequency segment , as shown in the following formula:

[0068]

[0069] Frequency importance value The larger the value, the more important the frequency segment f is;

[0070] Step 4.2: Calculate the frequency attention weight of the two-dimensional time-frequency spectrogram based on the frequency importance value of the two-dimensional time-frequency spectrogram;

[0071] Use a multi-layer perceptron network and calculate the weight vector of the frequency importance value of each frequency segment through the Softmax normalization layer , as shown in the following formula:

[0072]

[0073] Among them, K and b are learnable parameters, and ;

[0074] Step 5: Calculate the weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram based on the frequency attention weight of the two-dimensional time-frequency spectrogram and the one-dimensional feature vector of each sub-block;

[0075] The weight vector A of the frequency importance value of each frequency segment f Sequentially compare the one-dimensional feature vector of each sub-block Multiply element by element and weight the two-dimensional time-frequency spectrogram according to the frequency importance to obtain the weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram , as shown in the following formula:

[0076]

[0077] Step 6: Concatenate the weighted feature vectors of each sub-block in sequence to obtain the machine voiceprint feature matrix;

[0078] The feature vector of each sub-block Splice them in sequence to get the machine voiceprint feature matrix , as shown in the following formula:

[0079]

[0080] The machine voiceprint feature matrix V takes into account the information of key time frames and frequency frames at the same time, which can achieve a more robust and generalized feature representation of the machine voiceprint characteristics, and then serve as the input of subsequent fault detection, identification and diagnosis modules to improve their diagnostic accuracy and robustness.

[0081] This embodiment proposes a machine voiceprint feature modeling method, which divides the time-frequency spectrogram into multiple local blocks in the time dimension and performs weighted processing on each block in turn to enhance the modeling ability of key time series areas, and solve the problem of coarse feature representation granularity and neglect of dynamic time dimension information caused by global pooling in traditional methods; introduces a frequency importance evaluation mechanism and a sorting pooling strategy, and performs significant weighting on the frequency dimension features in each time block, thereby highlighting the diagnostic value of key frequency bands and improving the discrimination of voiceprint representations under different machine types and different working conditions; designs a structured local representation modeling process Based on the original two-dimensional time-frequency spectrogram, it is segmented and modeled through a unified block, adaptive weighting and aggregation method, avoiding the loss of key information caused by global sorting average pooling, and achieving a more targeted, robust and generalized voiceprint feature matrix representation, effectively enhancing the local sensitivity and fine-grained recognition ability of the model; with good compatibility and versatility, the audio voiceprint representation modeling method proposed in this invention is lightweight and highly scalable, and can be efficiently integrated into the existing industrial sound signal diagnosis process as a feature extraction process, and is suitable for voiceprint modeling tasks under various equipment types and various operating conditions;

[0082] Compared with the existing audio feature extraction method, the machine voiceprint feature modeling method of this embodiment can adaptively adjust and optimize hyperparameters according to the machine type in the actual industrial scenario. It can model a more robust voiceprint feature representation as the input of the diagnostic model, improve the accuracy and interpretability of the diagnostic results, and has good practical promotion prospects.

[0083] Example 2

[0084] This embodiment proposes a machine voiceprint feature modeling system, such as Figure 2 As shown, it includes: industrial sound signal data acquisition module, sound preprocessing and denoising module, and machine voiceprint characterization and modeling module.

[0085] The industrial sound signal data acquisition module is used to collect the sound signals of industrial equipment operation, store the collected industrial equipment operation sound signals in the form of audio signals, and transmit them to the sound preprocessing and denoising module;

[0086] The sound preprocessing and denoising module is used to preprocess the collected sound signals of industrial equipment operation. It uses Mel transform to convert the one-dimensional audio signal into a two-dimensional time-frequency spectrogram, and then transmits the two-dimensional time-frequency spectrogram to the machine voiceprint representation and modeling module.

[0087] The machine voiceprint representation and modeling module is used to process the two-dimensional time-frequency spectrogram, divide the two-dimensional time-frequency spectrogram into several sub-blocks in the time dimension, and extract the one-dimensional feature vector of each sub-block; at the same time, the weighted feature vector of each sub-block is calculated based on the frequency attention weight; finally, the weighted feature vectors of each sub-block are spliced ​​in sequence to obtain the machine voiceprint feature matrix.

[0088] Example 3

[0089] Based on the machine voiceprint feature modeling method, this embodiment proposes a method for voiceprint feature modeling and fault diagnosis in an abnormal transformer state. The voiceprint feature modeling process of the sound signal in the abnormal transformer operation state adopts the machine voiceprint feature modeling method of Example 1, including the following steps:

[0090] S1: collecting sound signals of the transformer in an abnormal operating state, preprocessing the collected sound signals of the transformer in an abnormal operating state, and obtaining a two-dimensional time-frequency spectrum X of the sound signals in the transformer in an abnormal operating state;

[0091] The preprocessing of the sound signal under abnormal operation state of the transformer includes: pre-emphasis, framing, windowing, denoising and short-time Fourier transform operations on the sound signal under abnormal operation state of the transformer, and using Mel transform to convert the one-dimensional audio signal into a two-dimensional time-frequency spectrogram , where F is the frequency dimension and T is the time dimension;

[0092] S2: Divide the two-dimensional time-frequency spectrum X of the sound signal under the abnormal operation state of the transformer into several non-overlapping sub-blocks;

[0093] The two-dimensional time-frequency spectrogram X of the sound signal under abnormal operation of the transformer is evenly divided into N non-overlapping sub-blocks of length W along the time dimension T. , , where the time length W=T / N;

[0094] S3: Sequentially sort the time frames in each sub-block of the two-dimensional time-frequency spectrogram of the sound signal under the abnormal operation state of the transformer based on the importance of the time dimension, and then convert each sub-block of the two-dimensional time-frequency spectrogram of the sound signal under the abnormal operation state of the transformer into a one-dimensional vector. Weighted pooling is performed on the sorted sub-blocks to obtain a one-dimensional feature vector for each sub-block;

[0095] S3.1: For each sub-block of the two-dimensional time-frequency spectrogram X of the sound signal under the abnormal operation state of the transformer, the energy of the frequency frame of each sub-block is sorted according to the time dimension. The two-dimensional time-frequency spectrogram after sorting is ;

[0096] S3.2: Using Learnable Pooling Vectors For the sorted sub-blocks Perform weighted pooling to obtain a one-dimensional feature vector that characterizes the voiceprint characteristics of each sub-block in a local time period. , as shown in the following formula:

[0097]

[0098]

[0099] Where T represents the matrix transpose calculation, r∈[0,1] is a learnable parameter, The sum of learnable parameters r;

[0100] S4: Calculate the frequency attention weight of the two-dimensional time-frequency spectrogram of the sound signal under the abnormal operation state of the transformer;

[0101] S4.1: Calculate the frequency importance value of the two-dimensional time-frequency spectrogram of the sound signal under abnormal operating conditions of the transformer;

[0102] Calculate the global average value of each frequency segment f∈F in the time dimension T of the two-dimensional time-frequency spectrogram X of the sound signal under abnormal operation of the transformer as the frequency importance value to measure the overall energy level of each frequency segment , as shown in the following formula:

[0103]

[0104] Frequency importance value The larger the value, the more important the frequency segment f is;

[0105] S4.2: Calculate the frequency attention weight of the two-dimensional time-frequency spectrogram of the sound signal under the abnormal operation state of the transformer based on the frequency importance value of the two-dimensional time-frequency spectrogram of the sound signal under the abnormal operation state of the transformer;

[0106] Use a multi-layer perceptron network and calculate the weight vector of the frequency importance value of each frequency segment through the Softmax normalization layer , as shown in the following formula:

[0107]

[0108] Among them, K and b are learnable parameters, and ;

[0109] S5: Calculate the weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram of the sound signal under the abnormal operation state of the transformer based on the frequency attention weight and the one-dimensional feature vector of each sub-block;

[0110] The weight vector A of the frequency importance value of each frequency segment of the sound signal under the abnormal operation state of the transformer is f Sequentially compare the one-dimensional feature vector of each sub-block Multiply element by element and weight the two-dimensional time-frequency spectrogram of the sound signal under abnormal operation of the transformer according to the frequency importance to obtain the weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram of the sound signal under abnormal operation of the transformer , as shown in the following formula:

[0111]

[0112] S6: The weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram of the sound signal under abnormal operation of the transformer By splicing them in sequence, we can get the machine voiceprint feature matrix of the sound signal under abnormal operation state of the transformer. , as shown in the following formula:

[0113]

[0114] S7: The machine voiceprint feature matrix of the sound signal under abnormal operation state of the transformer Input the fault diagnosis model to obtain the transformer fault diagnosis results;

[0115] In this embodiment, the SVM model, BP model and LSTM model are used as fault diagnosis models respectively. Figure 3This is a schematic diagram of the intermediate process results of the acoustic signal characterization modeling part of the transformer in an abnormal state using the proposed machine voiceprint feature modeling method in this embodiment, wherein (a) is the original spectrogram of the sound signal in the transformer in an abnormal operating state, (b) is the sorted two-dimensional time-frequency spectrum of the sound signal in the transformer in an abnormal operating state, (c) is the pooling vector, and (d) is the machine voiceprint feature matrix of the sound signal in the transformer in an abnormal operating state. In order to verify the effectiveness of the present invention in modeling the acoustic signal of the industrial machine operating state in practice, this embodiment selected an audio signal sample in an abnormal state of the transformer, and sequentially performed original spectrogram extraction, time dimension sorting, construction of pooling vectors, and output of the corresponding time-frequency feature representation matrix, corresponding to Figure 3 The results of (a), (b), (c) and (d) clearly show that the present invention can further compress redundant information while retaining the main time and frequency structures, thus achieving effective modeling of voiceprint representation. Figure 3 (d) The modeled voiceprint feature matrix shows that the proposed method can highlight the main energy areas of abnormal acoustic signals and effectively suppress invalid frequency bands and time frames, thereby providing more robust and discriminative input features for subsequent diagnostic models. Therefore, this also verifies the practicality and effectiveness of the modeling method in industrial scenarios.

[0116] Table 1 compares the recognition accuracy of three fault diagnosis models for different transformer faults, before and after using the machine voiceprint feature modeling method proposed in this embodiment. The three fault diagnosis models are SVM, BP, and LSTM. To verify the effectiveness and practicality of the present invention in modeling the operating acoustic signals of industrial equipment, this embodiment uses a dataset of transformer operating status audio data collected from a traditional method using raw spectrogram feature extraction, as well as features extracted using the voiceprint representation modeling method proposed in this embodiment. To ensure fairness in the experiment, the training set and test set ratio was 4:1. Three fault diagnosis models commonly used in practical applications were selected as the fault diagnosis models, and their default parameter settings were maintained. The recognition accuracy results of the three fault diagnosis models using the traditional method and the present invention are shown in Table 1. As can be seen from the results in Table 1, the recognition accuracy of all three diagnostic models improved after feature modeling using the present invention's method compared to the basic feature extraction method, demonstrating that the feature representation extracted by the present invention helps the model learn key information for fault identification.

[0117] Table 1 Comparison of recognition results of three fault diagnosis models before and after the voiceprint feature modeling method of this embodiment

[0118]

[0119] Through the above Figure 3 The experimental results in Table 1 verify the effectiveness and universality of the present invention in improving the voiceprint feature modeling capability of industrial sound signals. It can be applied as a universal voiceprint feature extraction and modeling solution to fault detection and diagnosis tasks of industrial equipment, significantly improving the recognition accuracy and robustness of the model.

[0120] Example 4

[0121] This embodiment proposes an electronic device, comprising: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the machine voiceprint feature modeling method.

[0122] The electronic device can be a mobile phone, computer, or tablet computer, and includes a memory and a processor. The memory stores a computer program that, when executed by the processor, implements the machine voiceprint feature modeling method described in the embodiments. It is understood that the electronic device may also include an input / output (I / O) interface and a communication component.

[0123] The processor is used to execute all or part of the steps in the machine voiceprint feature modeling method described in the above embodiment. The memory is used to store various types of data, such as instructions for any application or method in the electronic device, as well as application-related data.

[0124] The processor can be an application specific integrated circuit (ASIC), a digital signal processor (DSP), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor or other electronic components, and is used to execute the machine voiceprint feature modeling method described in the above embodiments.

[0125] Example 5

[0126] This embodiment provides a computer-readable storage medium storing executable instructions. When the instructions are executed, if they are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0127] The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the machine voiceprint feature modeling method described in each embodiment of the present application.

[0128] The aforementioned storage media include: flash memory, hard disk, multimedia card, card-type memory (for example, SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR abbreviation, memory data register) memory, etc.), random access memory (RAM, Random Access Memory), static random access memory (SRAM, Static Random-Access Memory), read-only memory (ROM, Read-Only Memory), electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read Only Memory), programmable read-only memory (PROM, Programmable Read-only Memory), magnetic memory, disk, optical disk, server, APP (Application, abbreviation of application software) application store and other media that can store program verification codes, on which a computer program is stored. When the computer program is executed by the processor, it can implement the various steps of the above-mentioned machine voiceprint feature modeling method.

[0129] Example 6

[0130] This embodiment provides a computer program product, including a computer program or instructions, which implements the machine voiceprint feature modeling method when executed by a processor.

[0131] Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a computer program product.

[0132] The various embodiments in this application are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0133] The scope of protection of this application is not limited to the above-described embodiments. Obviously, those skilled in the art may make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of this disclosure and its equivalents, the disclosure is intended to include such modifications and variations.

Claims

1. A machine voiceprint feature modeling method, characterized in that: The following steps are involved: Step 1: Collect the sound signals of industrial equipment operation, pre-process the sound signals of industrial equipment operation, and generate a two-dimensional time-frequency spectrogram; Step 2: Divide the two-dimensional time-frequency spectrogram into several non-overlapping sub-blocks; Step 3: Sort the time frames in each sub-block of the two-dimensional time-frequency spectrogram by their importance based on the time dimension, and then convert each sub-block of the two-dimensional time-frequency spectrogram into a one-dimensional vector. Perform weighted pooling on the sorted sub-blocks to obtain a one-dimensional feature vector for each sub-block. Step 4: Calculate the frequency attention weight of the two-dimensional time-frequency spectrogram; Step 5: Calculate the weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram based on the frequency attention weight of the two-dimensional time-frequency spectrogram and the one-dimensional feature vector of each sub-block; Step 6: Concatenate the weighted feature vectors of each sub-block in sequence to obtain the machine voiceprint feature matrix.

2. The machine voiceprint feature modeling method according to claim 1, characterized in that: The pre-processing of the industrial equipment operation sound signal in step 1 includes: pre-emphasis, framing, windowing, denoising and short-time Fourier transform operations on the industrial equipment operation sound signal, and using Mel transform to convert the one-dimensional audio signal into a two-dimensional time-frequency spectrogram. , where F is the frequency dimension and T is the time dimension.

3. The machine voiceprint feature modeling method according to claim 1, characterized in that: The specific method of step 2 is: Divide the two-dimensional time-frequency spectrogram X into N non-overlapping sub-blocks of length W along the time dimension T. , , where the time length W=T / N.

4. The machine voiceprint feature modeling method according to claim 1, characterized in that: The step 3 comprises: Step 3.1: For each sub-block of the two-dimensional time-frequency spectrogram, sort the energy of the frequency frame of each sub-block according to the time dimension. After sorting, the two-dimensional time-frequency spectrogram is ; Step 3.2: Use learnable pooling vectors For the sorted sub-blocks Perform weighted pooling to obtain a one-dimensional feature vector that characterizes the voiceprint characteristics of each sub-block in a local time period. , as shown in the following formula: ; ; Where T represents the matrix transpose calculation, r∈[0,1] is a learnable parameter, is the sum of the learnable parameters r.

5. The machine voiceprint feature modeling method according to claim 1, characterized in that: The step 4 comprises: Step 4.1: Calculate the frequency importance value of the two-dimensional time-frequency spectrogram; Step 4.2: Based on the frequency importance value of the two-dimensional time-frequency spectrogram, calculate the frequency attention weight of the two-dimensional time-frequency spectrogram.

6. The machine voiceprint feature modeling method according to claim 5, characterized in that: The specific method of step 4.1 is: Calculate the global average value of each frequency segment f∈F in the time dimension T of the two-dimensional time-frequency spectrogram X as the frequency importance value to measure the overall energy level of the frequency segment , as shown in the following formula: ; Frequency importance value The larger the value, the more important the frequency segment f is.

7. The machine voiceprint feature modeling method according to claim 6, characterized in that: The specific method of step 4.2 is: Use a multi-layer perceptron network and calculate the weight vector of the frequency importance value of each frequency segment through the Softmax normalization layer , as shown in the following formula: ; Among them, K and b are learnable parameters, and .

8. The machine voiceprint feature modeling method according to claim 1, characterized in that: The specific method of step 5 is: The weight vector A of the frequency importance value of each frequency segment f Sequentially with the one-dimensional feature vector of each sub-block Multiply element by element and weight the two-dimensional time-frequency spectrogram according to the frequency importance to obtain the weighted feature vector of each sub-block of the two-dimensional time-frequency spectrogram , as shown in the following formula: 。 9. The machine voiceprint feature modeling method according to claim 1, characterized in that: The specific method of step 6 is: The feature vector of each sub-block Splice them in sequence to get the machine voiceprint feature matrix , as shown in the following formula: 。 10. A machine voiceprint feature modeling system, which performs machine voiceprint feature modeling based on the method of claim 1, characterized in that: include: Industrial sound signal data acquisition module, sound preprocessing and denoising module, machine voiceprint characterization and modeling module; The industrial sound signal data acquisition module is used to collect the sound signals of industrial equipment operation, store the collected industrial equipment operation sound signals in the form of audio signals, and transmit them to the sound preprocessing and denoising module; The sound preprocessing and denoising module is used to preprocess the collected sound signals of industrial equipment operation. It uses Mel transform to convert the one-dimensional audio signal into a two-dimensional time-frequency spectrogram, and then transmits the two-dimensional time-frequency spectrogram to the machine voiceprint representation and modeling module. The machine voiceprint representation and modeling module is used to process the two-dimensional time-frequency spectrogram, divide the two-dimensional time-frequency spectrogram into several sub-blocks in the time dimension, and extract the one-dimensional feature vector of each sub-block; at the same time, the weighted feature vector of each sub-block is calculated based on the frequency attention weight; finally, the weighted feature vectors of each sub-block are spliced ​​in sequence to obtain the machine voiceprint feature matrix.

Citation Information

Patent Citations

  • Transformer voiceprint recognition method based on channel attention mechanism and AH-Softmax

    CN117497003A

  • Method for training transformer fault detection model, fault diagnosis method, and related device

    US20250370068A1

  • Audio-based user state identification method and apparatus, and electronic device and storage medium

    WO2021189903A1