Power equipment monitoring method and system based on multi-modal information cooperation

Through the power equipment monitoring method with multimodal information collaborative, the time-frequency analysis and coupling of soundprints and vibration information is used to generate a combined sound and vibration feature map, which solves the problem of incomplete state grasp in traditional monitoring methods and achieves comprehensive and accurate monitoring of power equipment.

CN120260609APending Publication Date: 2025-07-04CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510426254.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional power equipment monitoring methods are based on single mode information, making it difficult to fully and accurately grasp the equipment status, resulting in the inability to effectively monitor the operating status of power equipment.

Method used

By using multimodal information collaboration method, the soundprint and vibration information of the power equipment are obtained, time-frequency analysis is carried out, the coupling coefficients between the soundprint characteristics and vibration characteristics are calculated, and the time-frequency compounding is performed to generate a sound-vibration joint feature map to detect the equipment status.

Benefits of technology

It realizes comprehensive and accurate monitoring of the status of power equipment, can promptly detect potential problems, and ensure the stable operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260609A_ABST
    Figure CN120260609A_ABST
Patent Text Reader

Abstract

The invention discloses a power equipment monitoring method and system based on multi-modal information collaboration, and the method comprises the steps: firstly obtaining a to-be-monitored voiceprint and to-be-analyzed vibration information corresponding to the voiceprint, obtaining a voiceprint feature and a vibration feature through time-frequency analysis, calculating the coupling coefficient of the voiceprint feature and the vibration feature in each time-frequency unit to reflect the correlation strength, and carrying out the time-frequency analysis of the voiceprint feature and the vibration feature; and finally, detecting the state of at least one piece of to-be-detected power equipment in the to-be-monitored voiceprint segment according to the atlas and voiceprint features, thereby realizing multi-modal information cooperative monitoring of the power equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment maintenance, and more specifically, to a power equipment monitoring method and system based on multi-modal information collaboration. Background Art

[0002] With the continuous expansion of the scale of the power system, the reliable operation of power equipment is crucial. Traditional monitoring methods often rely on single-modal information and are difficult to comprehensively and accurately grasp the equipment status. Summary of the Invention

[0003] The purpose of the present invention is to provide a power equipment monitoring method and system based on multi-modal information collaboration.

[0004] In a first aspect, an embodiment of the present invention provides a power equipment monitoring method based on multi-modal information collaboration, including:

[0005] Obtaining a to-be-monitored voiceprint and vibration information to be analyzed corresponding to the to-be-monitored voiceprint, where the to-be-monitored voiceprint includes a plurality of to-be-monitored voiceprint segments;

[0006] Performing time-frequency analysis on the to-be-monitored voiceprint segments to obtain voiceprint features at multiple resolution levels, and performing time-frequency analysis on the vibration information to be analyzed to obtain vibration features;

[0007] Calculating a coupling coefficient between the voiceprint features and the vibration features in each time-frequency unit, where the coupling coefficient reflects the correlation strength between the voiceprint features and the vibration features in the same time-frequency unit;

[0008] Based on the coupling coefficient, performing time-frequency compounding on the voiceprint features and the vibration features to obtain a voice-vibration combined feature map;

[0009] Detecting at least one to-be-detected power equipment state in the to-be-monitored voiceprint segments according to the voice-vibration combined feature map and the voiceprint features.

[0010] In a second aspect, an embodiment of the present invention provides a power equipment monitoring system based on multi-modal information collaboration, characterized by including:

[0011] An initial unit for obtaining a to-be-monitored voiceprint and vibration information to be analyzed corresponding to the to-be-monitored voiceprint, where the to-be-monitored voiceprint includes a plurality of to-be-monitored voiceprint segments;

[0012] An analysis unit for performing time-frequency analysis on each of the plurality of to-be-monitored voiceprint segments to obtain voiceprint features at multiple resolution levels, and performing time-frequency analysis on the vibration information to be analyzed to obtain vibration features;

[0013] A calculation unit for calculating the coupling coefficient of the voiceprint feature and the vibration feature in each time-frequency unit, where the coupling coefficient reflects the correlation strength between the voiceprint feature and the vibration feature in the same time-frequency unit;

[0014] A composite unit for performing time-frequency composition on the voiceprint feature and the vibration feature based on the coupling coefficient to obtain a voice-vibration combined feature map;

[0015] A detection unit for detecting at least one power equipment state to be detected in the voiceprint segment to be monitored according to the voice-vibration combined feature map and the voiceprint feature.

[0016] In a third aspect, an embodiment of the present invention provides a server system, including a server, and the server is used to execute the method described in the first aspect.

[0017] In a fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the method described above is implemented.

[0018] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a power equipment monitoring method and system based on multi-modal information collaboration disclosed by the present invention, by obtaining the voiceprint to be monitored and its corresponding vibration information to be analyzed, respectively obtaining the voiceprint feature and the vibration feature through time-frequency analysis, calculating the coupling coefficient of the two in each time-frequency unit to reflect the correlation strength, then performing time-frequency composition based on this to obtain a voice-vibration combined feature map, and finally detecting at least one power equipment state to be detected in the voiceprint segment to be monitored according to the map and the voiceprint feature, realizing the collaborative monitoring of power equipment by multi-modal information. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic flowchart of the steps of the power equipment monitoring method based on multi-modal information collaboration provided by the embodiment of the present invention;

[0021] Figure 2 It is a schematic block diagram of the structure of the computer device provided by the embodiment of the present invention. Detailed Embodiments

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0023] The following will, with reference to the accompanying drawings, provide a detailed description of the specific implementation manners of the present invention.

[0024] To solve the technical problems in the foregoing background art, Figure 1 The following is a schematic flowchart of a power equipment monitoring method based on multimodal information collaboration provided for an embodiment of the present disclosure. The following will provide a detailed introduction to the power equipment monitoring method based on multimodal information collaboration.

[0025] Step S201: Obtain the to-be-monitored voiceprint and the vibration information to be analyzed corresponding to the to-be-monitored voiceprint, where the to-be-monitored voiceprint includes multiple to-be-monitored voiceprint segments;

[0026] Step S202: Perform time-frequency analysis on the to-be-monitored voiceprint segments to obtain voiceprint features at multiple resolution levels, and perform time-frequency analysis on the vibration information to be analyzed to obtain vibration features;

[0027] Step S203: Calculate the coupling coefficient between the voiceprint features and the vibration features in each time-frequency unit, where the coupling coefficient reflects the correlation strength between the voiceprint features and the vibration features in the same time-frequency unit;

[0028] Step S204: Based on the coupling coefficient, perform time-frequency compounding on the voiceprint features and the vibration features to obtain a voice-vibration combined feature map;

[0029] Step S205: Detect at least one to-be-detected power equipment state in the to-be-monitored voiceprint segments according to the voice-vibration combined feature map and the voiceprint features.

[0030] In an embodiment of the present invention, by way of example, in a large-scale substation, the server is responsible for data acquisition and processing of the entire power equipment monitoring system. There are numerous power equipment in the substation, such as transformers, switchgear, etc. The server obtains the to-be-monitored voiceprints through high-precision microphone arrays installed near each power equipment. These voiceprints contain various sound information emitted during the operation of the equipment and are segmented into multiple to-be-monitored voiceprint segments. For example, each segment may correspond to a small time interval of the equipment operation, facilitating subsequent analysis and processing. At the same time, the server also obtains the to-be-analyzed vibration information corresponding to the to-be-monitored voiceprints through vibration sensors installed at key parts of the equipment. For example, for a transformer, the vibration sensors are installed at its iron core, winding and other parts to accurately capture the vibration conditions generated during the operation of the equipment. These vibration information are correlated with the voiceprint information, providing multi-faceted data support for comprehensively understanding the equipment status. After receiving the to-be-monitored voiceprint segments and the to-be-analyzed vibration information, the server starts time-frequency analysis. For the to-be-monitored voiceprint segments, the server adopts advanced time-frequency analysis algorithms, such as the short-time Fourier transform (STFT). Taking a transformer as an example, during normal operation, the sound it emits has certain frequency characteristics and time distribution rules. The server converts the voiceprint segments from the time domain to the time-frequency domain through STFT, obtaining the energy distribution at different time points and frequency points. After corresponding processing, multi-resolution level voiceprint characteristics are finally obtained. These voiceprint characteristics can more finely depict the characteristics of the voiceprint at different time and frequency scales, just like drawing a multi-dimensional portrait of the voiceprint information, clearly presenting various detailed changes in the sound. Similarly, for the to-be-analyzed vibration information, the server also uses a similar time-frequency analysis method. Taking the switchgear as an example, when abnormal conditions such as loose electrical connection components occur inside the switchgear, vibrations with specific frequencies and amplitudes will be generated. The server converts the vibration information to the time-frequency domain through time-frequency analysis, obtaining vibration characteristics. These vibration characteristics can accurately reflect the frequency, amplitude of the vibration and its changes at different times, providing an important basis for subsequent judgment of whether there are faults in the equipment. After obtaining the voiceprint characteristics and vibration characteristics, the server starts to calculate their coupling coefficients in each time-frequency unit. Suppose on a large transformer in the substation, the server has obtained its voiceprint characteristics and vibration characteristics. The voiceprint characteristics can contain information in multiple frequency band dimensions. For example, the low-frequency band may correspond to the sound generated by magnetostriction of the transformer iron core, and the high-frequency band may be related to the electromagnetic vibration of the winding. The vibration characteristics also have different frequency band dimensions. For example, some low-frequency vibrations may originate from the mechanical structure vibration of the equipment, and high-frequency vibrations may be related to the high-frequency vibrations of electrical components. The server first identifies the number of frequency band dimensions in the voiceprint characteristics to obtain the voiceprint frequency band number, and the number of frequency band dimensions in the vibration characteristics to obtain the vibration frequency band number. Then, based on these two numbers, the target frequency band dimension number of the time-frequency representation domain corresponding to each resolution level is determined.Next, the vibrator modes are superimposed in the time domain to obtain the superimposed vibration characteristics, and the dimensionality of the voiceprint characteristics and the superimposed vibration characteristics are respectively optimized to the target frequency band dimensionality to obtain the aligned voiceprint characteristics and the aligned vibration characteristics. For example, within a certain time-frequency unit, the server extracts the voiceprint sub-characteristics corresponding to each time-frequency unit from the aligned voiceprint characteristics, and further extracts the voiceprint frequency band characteristics. At the same time, in the aligned vibration characteristics, the characteristics corresponding to the frequency band dimensionality corresponding to the voiceprint frequency band characteristics are located to obtain the vibration frequency band characteristics. Then, the voiceprint frequency band characteristics and the vibration frequency band characteristics are coupled to obtain the coupled channel responses corresponding to each frequency band dimensionality. The coupled channel responses within each time-frequency unit are integrated to obtain the joint coupled response, and then the voiceprint sub-characteristics and the aligned vibration characteristics are coupled to obtain the reference coupled response. Finally, the energy ratio between the reference coupled response and the joint coupled response is calculated to obtain the coupling coefficient corresponding to this time-frequency unit. This coupling coefficient can accurately reflect the correlation strength between the voiceprint characteristics and the vibration characteristics within this specific time-frequency unit, just like building a bridge between the information of these two different modes to measure their closeness. Based on the calculated coupling coefficient, the server performs time-frequency compounding on the voiceprint characteristics and the vibration characteristics to obtain the acoustic-vibration joint feature map. Taking a group of switchgear in a substation as an example, the server performs energy ratio matching on the voiceprint sub-characteristics based on the previously calculated coupling coefficient. For example, for a time-frequency unit with a relatively high coupling coefficient, it indicates that the correlation between the voiceprint characteristics and the vibration characteristics within this unit is strong, and the server will correspondingly increase the energy weight of the voiceprint sub-characteristics corresponding to this time-frequency unit during compounding. Through such energy ratio matching, the reference acoustic-vibration joint feature maps corresponding to each time-frequency unit are obtained. Then, the server performs time-frequency compounding on the reference acoustic-vibration joint feature maps at the same resolution level. Suppose the server uses a pre-trained deep learning model to complete this compounding process, and this deep learning model is trained based on the acoustic-vibration joint feature map data of power equipment in normal and abnormal states with annotations. The server inputs the reference acoustic-vibration joint feature maps at the same resolution level as the input data into this model, and the model performs feature extraction and fusion processing on the input data through its convolutional layer, pooling layer, and fully connected layer structures, and finally outputs the acoustic-vibration joint feature maps at each resolution level. These acoustic-vibration joint feature maps integrate the information of both the voiceprint and vibration modes and can more comprehensively and accurately reflect the operating state of power equipment. The server uses the obtained acoustic-vibration joint feature maps and voiceprint characteristics to detect at least one power equipment state to be detected in the voiceprint segment to be monitored. Taking a transformer in a substation as an example, the server first performs hierarchical optimization on the voiceprint characteristics based on the resolution level of the voiceprint characteristics. For example, for some resolution levels that can more clearly present the voiceprint characteristics related to the vibration of the transformer core, the server will preferentially select these levels.Then, based on the preferred hierarchical sequence, the voiceprint features corresponding to the target resolution level are located in the voiceprint features to obtain dynamic voiceprint features. At the same time, the server locates the corresponding acoustic-vibration joint feature map at the target resolution level in the acoustic-vibration joint feature map to obtain a dynamic acoustic-vibration joint feature map. Next, the server performs eigen-coupling analysis on the device state components to obtain a transient device state representation. For example, for a transformer, its magnetization state of the iron core, temperature state of the winding, etc. may be analyzed as device state components, and the transient representation of these states at a certain moment is obtained through eigen-coupling analysis. Then, the server performs cross-modal coupling on the transient device state representation, the dynamic acoustic-vibration joint feature map, and the dynamic voiceprint features to obtain a dynamic device state representation. After that, the server calibrates the dynamic device state representation to the time-frequency representation domain of the preset device state representation to obtain an intermediate device state representation, and uses it as the preset device state representation. At the same time, the server sets preset convergence criteria, such as the threshold of the device state representation difference rate and the threshold of the phase coherence coefficient. The server recursively triggers the above process from performing eigen-coupling analysis on the device state components to calibrating the dynamic device state representation to the time-frequency representation domain of the preset device state representation to obtain an intermediate device state representation, until the preset convergence criteria are reached and the recursive process is terminated, obtaining an incrementally corrected device state representation, and using it as the final preset device state representation. Finally, the server recursively triggers the process of locating the voiceprint features corresponding to the target resolution level in the voiceprint features based on the preferred hierarchical sequence, until the recursive process is terminated when all voiceprint features are dynamic voiceprint features, obtaining a device state representation. Based on this device state representation, the server can accurately determine at least one state of the power device to be detected in the voiceprint segment to be monitored, such as determining whether there are faults such as loose iron core and overheating of the winding in the transformer, so as to realize the effective monitoring of the power device.

[0031] In the embodiment of the present invention, calculating the coupling coefficient between the voiceprint feature and the vibration feature in each time-frequency unit can be implemented through the following examples.

[0032] Calibrate the voiceprint feature and the vibration feature to the same time-frequency representation domain to obtain the aligned voiceprint features at each resolution level and the aligned vibration features corresponding to the aligned voiceprint features;

[0033] Extract the voiceprint sub-features corresponding to each time-frequency unit from the aligned voiceprint features, and calculate the coupling coefficient between the voiceprint sub-features and the aligned vibration features.

[0034] In an embodiment of the present invention, exemplarily, in the power equipment monitoring scenario of a substation, the server is responsible for processing data collected from various devices. For example, for a large transformer, the server obtains its corresponding voiceprint feature and vibration feature. The voiceprint feature is obtained by processing the sound information collected by a high-precision microphone array installed near the transformer. It contains information in different frequency band dimensions and reflects the characteristics of the sound emitted by the transformer during operation at different frequencies and times. The vibration feature is collected by vibration sensors installed at key parts of the transformer, such as the iron core and winding, and also has various frequency band dimension representations corresponding to the vibration conditions of different components. The first thing the server needs to do is to calibrate these two features into the same time-frequency representation domain. It will first identify the number of frequency band dimensions in the voiceprint feature. Suppose the voiceprint feature has 5 different frequency band dimensions, and at the same time, it also determines the number of frequency band dimensions of the vibration feature, such as 4 frequency band dimensions. Then, based on certain rules and algorithms, considering both situations, it determines the target frequency band dimension number of the time-frequency representation domain corresponding to each resolution level. Suppose it is determined to be 3. Next, the server will perform corresponding processing on the vibration feature. For example, it will perform time-domain superposition on its different vibration sub-modes to obtain the superimposed vibration feature, and then optimize and adjust the dimension numbers of the voiceprint feature and the superimposed vibration feature to the target frequency band dimension number respectively, so as to obtain the aligned voiceprint feature and the corresponding aligned vibration feature at each resolution level. It's like transforming these two pieces of information with originally different "specifications" into a form that can be compared and analyzed under the same "standard framework". After obtaining the aligned voiceprint feature and the aligned vibration feature, the server starts further analysis. Taking a certain operation period of the transformer as an example, in the aligned voiceprint feature at a specific resolution level, the server will extract the voiceprint sub-features corresponding to each time-frequency unit according to the set algorithms and rules. For example, in the case where a time period is 10 seconds and the frequency range from 100 Hz to 1000 Hz is divided into several time-frequency units, the server accurately extracts the unique voiceprint sub-features within each time-frequency unit. These voiceprint sub-features can more precisely reflect the characteristics of the transformer sound in this specific time-frequency interval. Then, the server will calculate the coupling coefficient between these extracted voiceprint sub-features and the aligned vibration feature. For each voiceprint sub-feature, the server will measure the degree of tight association between it and the aligned vibration feature within the same time-frequency unit. For example, when the iron core of the transformer has a slight looseness at a certain moment, it may cause an abnormal sound with a specific frequency in a certain specific time-frequency unit in the voiceprint sub-feature, and at the same time, there will be a corresponding abnormal vibration performance in the vibration feature. By calculating the coupling coefficient, the correlation intensity between this voiceprint and vibration can be accurately quantified, providing a key basis for accurately judging the operating state of the transformer in the future.

[0035] In an embodiment of the present invention, the vibration feature includes at least one vibration sub-mode. To calibrate the voiceprint feature and the vibration feature into the same time-frequency representation domain to obtain the aligned voiceprint features at each resolution level and the aligned vibration features corresponding to the aligned voiceprint features, the following example can be used for implementation.

[0036] Identify the number of frequency band dimensions in the voiceprint feature to obtain the voiceprint frequency band number, and identify the number of frequency band dimensions in the vibration feature to obtain the vibration frequency band number;

[0037] Based on the voiceprint frequency band number and the vibration frequency band number, determine the target frequency band dimension number of the time-frequency representation domain corresponding to each resolution level, and perform time-domain superposition on the vibration sub-modes to obtain the superimposed vibration feature;

[0038] Optimize the dimension numbers of the voiceprint feature and the superimposed vibration feature to the target frequency band dimension number respectively to obtain the aligned voiceprint features at each resolution level and the aligned vibration features corresponding to the aligned voiceprint features.

[0039] In an embodiment of the present invention, exemplarily, in a large power plant, a server is responsible for monitoring numerous power equipment. Taking one of the key generators as an example, the server obtains the voiceprint features to be monitored through a microphone array installed near the generator. These voiceprint features contain rich information about the sound emitted during the operation of the generator. The server carefully analyzes the voiceprint features and accurately identifies the number of frequency band dimensions therein. For example, through detailed data sorting and algorithm processing, it is determined that the number of voiceprint frequency bands is 8, and each frequency band dimension corresponds to a specific manifestation of the sound under different frequency ranges. At the same time, the server obtains the vibration features to be analyzed through vibration sensors installed at key parts of the generator (such as the rotor, stator, etc.). The vibration features include at least one vibration sub-mode, reflecting the vibration conditions of different components of the generator. The server also deeply analyzes the vibration features and identifies the number of frequency band dimensions therein. Suppose the detected number of vibration frequency bands is 6, and different frequency band dimensions correspond to relevant information about different vibration frequencies and amplitudes. Based on the identified 8 voiceprint frequency band numbers and 6 vibration frequency band numbers, the server determines the target frequency band dimension number of the time-frequency representation domain corresponding to each resolution level according to a set of rules that combine the operating characteristics of power equipment and data analysis requirements. For example, after comprehensive consideration, the target frequency band dimension number is determined to be 5. This number can facilitate subsequent unified analysis and processing while ensuring the effective capture of key equipment information. Then, the server processes the vibration sub-modes in the vibration features. Since the vibration conditions of the generator are relatively complex and there are multiple vibration sub-modes, the server superimposes these vibration sub-modes in the time domain in chronological order. For example, the high-frequency vibration sub-mode of the rotor and the low-frequency vibration sub-mode of the stator, etc., are superimposed and integrated on the time axis through a specific algorithm to obtain the superimposed vibration features, so that they can be better compared and analyzed with the voiceprint features in the same time-frequency representation domain. Subsequently, the server optimizes the dimension numbers of the voiceprint features and the superimposed vibration features to the target frequency band dimension number of 5 respectively. For the voiceprint features, through a series of data processing and conversion operations, some relatively minor frequency band dimension information is removed, and the key part that matches the target frequency band dimension number is retained, thereby obtaining the aligned voiceprint features at each resolution level. Similarly, similar optimization processing is performed on the superimposed vibration features, adjusting its data structure and frequency band dimension distribution so that its dimension number also becomes 5, and then obtaining the aligned vibration features corresponding to the aligned voiceprint features. In this way, the server successfully calibrates the voiceprint features and the vibration features to the same time-frequency representation domain, laying a foundation for accurately analyzing the operating state of the generator subsequently.

[0040] In an embodiment of the present invention, to calculate the coupling coefficient between the voiceprint sub-features and the aligned vibration features, the following example can be executed.

[0041] Extract the features in each frequency band dimension from the voiceprint sub-features to obtain the voiceprint frequency band features;

[0042] Align the features in the corresponding frequency band dimension of the voiceprint frequency band features in the alignment vibration features to obtain the vibration frequency band features;

[0043] Based on the voiceprint frequency band features and the vibration frequency band features, determine the coupling coefficient between the voiceprint sub-features and the alignment vibration features.

[0044] In an embodiment of the present invention, exemplarily, in the power equipment monitoring work of a substation, the server is processing the data of a large transformer. Previously, the acoustic sub-features corresponding to each time-frequency unit have been extracted from the acoustic fingerprint information generated during the operation of the transformer. Now, the server conducts further analysis on each acoustic sub-feature. Taking the acoustic sub-feature within a specific time-frequency unit as an example, it records the specific manifestation of the sound of the transformer during this short period. The server delves into the interior of this acoustic sub-feature through a refined data processing algorithm and divides and extracts it according to different frequency band dimensions. For example, when the transformer is operating normally, the sound in the low-frequency band may be a relatively low-pitched humming sound emitted by the magnetostriction of the iron core, and in the high-frequency band may be a relatively sharp sound generated by the electromagnetic vibration of the winding. The server accurately extracts the features in each frequency band dimension from the acoustic sub-feature according to the frequency range, sorts and classifies them, thereby obtaining the acoustic frequency band features. These acoustic frequency band features can more clearly show the specific characteristics of the sound in different frequency band dimensions, just like further subdividing the "jigsaw puzzle" of the acoustic sub-feature according to the frequency band dimension to more accurately correlate and analyze with the vibration features. At the same time, the server has calibrated the vibration features and acoustic features of the transformer to the same time-frequency representation domain to obtain the aligned vibration features. For the acoustic frequency band features just extracted, the server needs to find the features in the corresponding frequency band dimension in the aligned vibration features. Still taking the transformer as an example, its vibration features reflect the vibration conditions of components such as the iron core and winding. When a certain frequency band dimension is determined in the acoustic frequency band features, such as the low-frequency band corresponding to the sound frequency band dimension related to the iron core, the server accurately aligns to the features reflecting the vibration conditions of the iron core in the same low-frequency band in the aligned vibration features. Through such meticulous alignment operations, the vibration frequency band features are obtained. This is like finding the corresponding parts in two different but related "information maps" according to the same frequency band dimension coordinates, enabling the accurate matching of the acoustic and vibration at the frequency band dimension level and preparing for the subsequent calculation of the coupling coefficient. After the server obtains the acoustic frequency band features and vibration frequency band features, it begins to determine the coupling coefficient between the acoustic sub-feature and the aligned vibration feature. Taking an instant during the operation of the transformer as an example, assume that within a certain time-frequency unit, the acoustic frequency band features show abnormal sharp sound features in the high-frequency band, which may imply some abnormal conditions in the winding. And in the corresponding vibration frequency band features, it is also found that the vibration amplitude of the winding in the high-frequency band has increased. The server comprehensively considers the performances of the acoustic frequency band features and vibration frequency band features in various aspects, such as the matching degree of frequencies and the variation relationship of amplitudes. Through the quantitative analysis and calculation of these factors, the coupling coefficient between the acoustic sub-feature and the aligned vibration feature within this time-frequency unit is finally determined. This coupling coefficient can accurately reflect the degree of tight correlation between the acoustic and vibration in this specific time-frequency unit, thereby providing a key basis for judging the operating state of the transformer, such as whether there are faults.

[0045] In an embodiment of the present invention, the coupling coefficient between the voiceprint sub - feature and the aligned vibration feature is determined based on the voiceprint frequency - band feature and the vibration frequency - band dimension, and the implementation can be carried out through the following examples.

[0046] Couple the voiceprint frequency - band feature with the vibration frequency - band feature to obtain the coupling channel response corresponding to each frequency - band dimension;

[0047] Integrate the coupling channel responses within each time - frequency unit to obtain the joint coupling response corresponding to each time - frequency unit;

[0048] Couple the voiceprint sub - feature with the aligned vibration feature to obtain a reference coupling response, and calculate the energy ratio between the reference coupling response and the joint coupling response to obtain the coupling coefficient corresponding to each time - frequency unit.

[0049] In an embodiment of the present invention, by way of example, in the scenario of monitoring power equipment in a large power plant, the server is responsible for analyzing the data of an important generator. Previously, the acoustic frequency band features have been extracted from the acoustic fingerprint information generated during the operation of the generator, and the vibration frequency band features have been obtained by aligning the vibration features. Now, the server operates for each frequency band dimension. For example, for the low-frequency band of the generator, the acoustic frequency band features may be manifested as relatively low-pitched sound features caused by the rotation of the rotor, and the vibration frequency band features reflect the vibration conditions of the rotor in the low-frequency band. The server couples the acoustic frequency band features and the vibration frequency band features in this frequency band dimension through a specific coupling algorithm. It is like deeply fusing the information describing the operation state of the same low-frequency band of the generator from different angles. After calculation, the coupling channel response corresponding to this frequency band dimension is obtained. Similarly, such coupling operations are performed on each frequency band dimension such as the middle-frequency band and high-frequency band of the generator, so as to obtain the coupling channel responses corresponding to each frequency band dimension. After the server obtains the coupling channel responses corresponding to each frequency band dimension, it starts to perform an integration operation on these coupling channel responses within each time-frequency unit. Taking a time-frequency unit corresponding to a certain period of the generator operation as an example, this time-frequency unit covers a certain time range and frequency range. The server accumulates and sums up the coupling channel responses obtained in different frequency band dimensions within this time-frequency unit according to the established integration rules. It is like integrating the "local responses" of each frequency band dimension within this specific time-frequency unit. After the integration operation, the joint coupling response corresponding to this time-frequency unit is obtained. Such operations are performed on each time-frequency unit during the operation of the generator, so as to obtain the joint coupling responses corresponding to each time-frequency unit. Then, the server couples the acoustic sub-features and the aligned vibration features to obtain a reference coupling response. For example, for the overall acoustic and vibration conditions of the generator at a certain moment, the acoustic sub-features and the aligned vibration features are deeply fused through a specific coupling method to obtain a reference coupling response reflecting the overall correlation degree. Then, the server calculates the energy ratio between the reference coupling response and the joint coupling response. Taking a specific time-frequency unit as an example, the energy value of the reference coupling response corresponding to this time-frequency unit is compared with the energy value of the joint coupling response, and through the ratio operation of the two energies, the coupling coefficient corresponding to this time-frequency unit is obtained. This coupling coefficient accurately reflects the degree of tight correlation between the acoustic sub-features and the aligned vibration features within this specific time-frequency unit, providing an important basis for judging the operation state of the generator.

[0050] In an embodiment of the present invention, based on the coupling coefficient, performing time-frequency compounding on the acoustic feature and the vibration feature to obtain an acoustic-vibration joint feature map can be implemented through the following examples.

[0051] Based on the coupling coefficient, perform energy ratio matching on the voiceprint sub-features to obtain a reference acoustic-vibration joint feature map corresponding to each time-frequency unit;

[0052] Perform time-frequency compounding on the reference acoustic-vibration joint feature maps at the same resolution level to obtain acoustic-vibration joint feature maps at each resolution level.

[0053] In an embodiment of the present invention, for example, in the scenario of monitoring power equipment in a substation, the server is processing data collected from a transformer. The coupling coefficients of the voiceprint features and the vibration features in each time-frequency unit have been calculated previously. Now, the server performs energy ratio matching operations on the voiceprint sub-features based on these coupling coefficients. Taking a certain time-frequency unit during a certain operation period of the transformer as an example, if the coupling coefficient corresponding to this time-frequency unit is high, it means that the association between the voiceprint feature and the vibration feature in this unit is very strong. The server will accordingly increase the energy weight of this voiceprint sub-feature in this time-frequency unit. For example, normally the energy value of this voiceprint sub-feature in the time-frequency unit is set to 50 (hypothetical quantization value), and due to the high coupling coefficient, the server will increase its energy value to 80 to highlight its importance. By performing such energy ratio matching operations on the voiceprint sub-features in each time-frequency unit according to their corresponding coupling coefficients, the server finally obtains a reference acoustic-vibration joint feature map corresponding to each time-frequency unit. These maps comprehensively incorporate the energy-adjusted voiceprint sub-features and the corresponding vibration feature information in different time-frequency units, and can more accurately reflect the comprehensive situation of the transformer operation state in this time-frequency unit. After the server obtains the reference acoustic-vibration joint feature maps corresponding to each time-frequency unit, it then needs to perform time-frequency compounding operations on these maps at the same resolution level. Still taking the transformer as an example, assume that the server has obtained multiple reference acoustic-vibration joint feature maps at a certain resolution level, and these maps correspond to different time-frequency units but are at the same resolution level. The server takes them as input data and processes them using a specific time-frequency compounding algorithm. It's like integrating and splicing these scattered but same-level "local maps" of the transformer operation state according to the rules of time and frequency. Through the operation and processing process, the server finally obtains acoustic-vibration joint feature maps at each resolution level. These maps completely present the overall picture of the transformer operation state that comprehensively incorporates the voiceprint and vibration modality information at each resolution level, providing a more comprehensive and accurate basis for subsequent accurately judging whether there are faults in the transformer, etc.

[0054] In an embodiment of the present invention, detecting at least one power equipment state to be detected in the voiceprint segment to be monitored according to the acoustic-vibration joint feature map and the voiceprint feature can be implemented through the following examples.

[0055] Based on the resolution levels of the voiceprint features, perform hierarchical optimization on the voiceprint features, and based on the optimized hierarchical sequence, align the voiceprint features at the target resolution level in the voiceprint features to obtain dynamic voiceprint features;

[0056] Align the acoustic-vibration combined feature map at the target resolution level in the acoustic-vibration combined feature map to obtain a dynamic acoustic-vibration combined feature map;

[0057] Extract the device state representation from the to-be-monitored voiceprint segment according to the dynamic acoustic-vibration combined feature map and the dynamic voiceprint features, and based on the device state representation, determine at least one to-be-detected power device state in the to-be-monitored voiceprint segment.

[0058] In the embodiments of the present invention, by way of example, in the power equipment monitoring work of a large power plant, the server is responsible for analyzing the operation status data of numerous generator sets. Taking one large generator set as an example, the server has obtained the corresponding to-be-monitored voiceprint segment and the voiceprint features obtained through a series of processes, and these voiceprint features have multiple resolution levels. The server will perform hierarchical optimization based on the resolution levels of the voiceprint features according to the operation characteristics of the generator set and past experience. For example, for the resolution level reflecting the sound characteristics related to the rotor vibration of the generator set, since it is crucial for judging the device state, the server will preferentially select it to form an optimized hierarchical sequence. Then, the server accurately aligns the voiceprint features at the target resolution level in the voiceprint features according to this optimized hierarchical sequence. Just like selecting the most suitable one from a stack of "voiceprint portraits" with different levels of detail according to specific criteria, so as to obtain dynamic voiceprint features. This dynamic voiceprint feature focuses more on the information level that is important for judging the device state. At the same time, the server also has the previously generated acoustic-vibration combined feature map. For the just-determined target resolution level, the server performs precise alignment operations in the acoustic-vibration combined feature map. Taking the acoustic-vibration combined feature map generated during the operation of the generator set as an example, the server quickly finds the part of the acoustic-vibration combined feature map corresponding to the target resolution level through specific algorithms and data marking, just like finding the detailed map of a specific area in a large "acoustic-vibration comprehensive map". Figure 1In this way, a dynamic acoustic-vibration joint feature map is obtained. After the server obtains the dynamic voiceprint features and the dynamic acoustic-vibration joint feature map, it begins to extract the device state representation from them. For a generator set, the server will analyze information such as vibration frequency and amplitude in the dynamic acoustic-vibration joint feature map, as well as changes in the frequency and timbre of the sound in the dynamic voiceprint features, and comprehensively use this information to extract the device state representation, such as determining whether there is imbalance in the rotor and whether the winding is overheated and other state information. Finally, based on these extracted device state representations, the server can clearly determine at least one power device state to be detected in the voiceprint segment to be monitored. For example, it is determined that the rotor of the generator set is currently in a slightly unbalanced state, thus realizing the precise monitoring of the operating state of the power device and the pre-judgment of faults.

[0059] In the embodiment of the present invention, the extracting of the device state representation from the voiceprint segment to be monitored according to the dynamic acoustic-vibration joint feature map and the dynamic voiceprint features can be implemented through the following examples.

[0060] Based on the dynamic acoustic-vibration joint feature map and the dynamic voiceprint features, perform incremental correction on the preset device state representation, and use the incrementally corrected device state representation as the preset device state representation;

[0061] Recursively trigger the process of aligning the voiceprint features to the voiceprint features at the target resolution level in the voiceprint features based on the preferred hierarchical sequence until the recursive process terminates when each voiceprint feature is the dynamic voiceprint feature, and obtain the device state representation.

[0062] In an embodiment of the present invention, exemplarily, in the scenario of monitoring power equipment in a substation, the server is responsible for analyzing the data of a large transformer. The server has obtained the dynamic acoustic-vibration combined feature map and the dynamic voiceprint feature of the transformer. The preset device state characterization may include initial set values or reference ranges of various indicators such as the state of the transformer core and the state of the winding. The server will perform incremental correction on these preset device state characterizations based on the currently obtained dynamic acoustic-vibration combined feature map and the dynamic voiceprint feature. For example, it is found from the dynamic acoustic-vibration combined feature map that the vibration frequency corresponding to the transformer winding in a certain period has a slight increase, and a slightly sharper sound than usual is heard in the dynamic voiceprint feature, which may imply an upward trend in the winding temperature. The server will adjust the indicator regarding the winding temperature state in the preset device state characterization according to this information, appropriately tighten the originally set normal temperature range, or slightly adjust the currently estimated winding temperature value, so as to complete the incremental correction of the preset device state characterization, and use the corrected result as the new preset device state characterization again. The server will then recursively trigger the process of aligning the voiceprint feature to the target resolution level in the voiceprint feature based on the preferred hierarchy sequence. Still taking the transformer as an example, initially, only some voiceprint features may obtain the dynamic voiceprint feature through hierarchical preference, and there are other voiceprint features that do not participate in this process. The server will again, according to the preferred hierarchy sequence, align the voiceprint feature to the target resolution level in the remaining voiceprint features, and repeat this process continuously. Each time this process is repeated, further incremental correction of the preset device state characterization will be performed based on the newly obtained dynamic voiceprint feature combined with the dynamic acoustic-vibration combined feature map. Repeating like this until all the voiceprint features have been transformed into dynamic voiceprint features through this process, at which point the recursive process terminates. The device state characterization finally obtained by the server is the result after integrating all relevant information in the entire monitoring process, through multiple corrections and improvements, and can accurately reflect the actual device state of the transformer during the period corresponding to the voiceprint segment to be monitored, such as accurately judging whether the core is loose, whether the winding is overheated, and other specific situations.

[0063] In an embodiment of the present invention, the incremental correction of the preset device state characterization based on the dynamic acoustic-vibration combined feature map and the dynamic voiceprint feature may be executed through the following examples.

[0064] Perform incremental correction on the preset device state characterization based on the dynamic acoustic-vibration combined feature map and the dynamic voiceprint feature to obtain an intermediate device state characterization, and use the intermediate device state characterization as the preset device state characterization;

[0065] Recursively trigger the process of incrementally correcting the preset device state representation based on the dynamic acoustic-vibration combined feature map and dynamic voiceprint features until the recursive process terminates when the preset convergence criterion is met, and obtain the incrementally corrected device state representation. The preset convergence criterion includes a device state representation difference rate threshold and a phase coherence coefficient threshold.

[0066] In an embodiment of the present invention, for example, in a large power plant, a server is responsible for monitoring the operating state of a key generator. The server has obtained the dynamic acoustic-vibration combined feature map and dynamic voiceprint features of the generator, as well as the initial preset device state representation, which covers various indicators such as the rotor balance state and stator winding temperature state of the generator. The server incrementally corrects the preset device state representation based on the current dynamic acoustic-vibration combined feature map and dynamic voiceprint features. For example, it is found from the dynamic acoustic-vibration combined feature map that the vibration amplitude of the generator rotor increases compared with the normal situation during a certain period, and abnormal changes in the sound at the corresponding frequency are also detected in the dynamic voiceprint features. Based on this, the server adjusts the indicator regarding the rotor balance state in the preset device state representation, perhaps slightly reducing the originally determined normal balance state to a critical state close to imbalance. At the same time, the indicator of the stator winding temperature state is also comprehensively judged based on the sound and vibration information to determine whether fine-tuning is required. After such processing, an intermediate device state representation is obtained, and then it is used as the new preset device state representation. The server will then recursively trigger the process of incrementally correcting the preset device state representation based on the dynamic acoustic-vibration combined feature map and dynamic voiceprint features. Each time of recursion, the current preset device state representation (i.e., the intermediate device state representation obtained in the previous round) is further incrementally corrected based on the new dynamic acoustic-vibration combined feature map and dynamic voiceprint features. For example, in subsequent monitoring, it is found that there are new changes in the vibration frequency of the generator stator winding, and there are also associated sound changes in the voiceprint features. The server will then adjust the relevant indicators such as the stator winding temperature and mechanical structure stability in the preset device state representation again. This recursive process will continue until the preset convergence criterion is met. For example, the device state representation difference rate threshold is set to 5%, and the phase coherence coefficient threshold is set to 0.8. When, after multiple rounds of correction, the difference rate of each indicator of the newly obtained device state representation compared with the previous round is less than 5%, and the phase coherence coefficient reaches above 0.8, the recursive process is terminated. At this time, the incrementally corrected device state representation obtained is the final result that integrates a large amount of monitoring information, is finely adjusted, and meets the convergence conditions, and can accurately reflect the actual operating state of the generator.

[0067] In an embodiment of the present invention, the preset device state representation includes at least one device state component. The incremental correction of the preset device state representation based on the dynamic acoustic-vibration joint feature map and the dynamic voiceprint feature to obtain the intermediate device state representation can be implemented through the following examples.

[0068] Perform eigen-coupling analysis on the device state component to obtain the transient device state representation;

[0069] Perform cross-modal coupling on the transient device state representation, the dynamic acoustic-vibration joint feature map, and the dynamic voiceprint feature to obtain the dynamic device state representation;

[0070] Calibrate the dynamic device state representation to the time-frequency representation domain of the preset device state representation to obtain the intermediate device state representation.

[0071] In an embodiment of the present invention, by way of example, in the scenario of monitoring the power equipment in a substation, the server is responsible for analyzing the operating status of a large transformer. The preset device status representation includes multiple device status components such as the core status and winding status of the transformer. The server performs eigen-coupling analysis on these device status components. Taking the core status component as an example, the server will deeply analyze the core vibration-related information obtained from the dynamic acoustic-vibration joint feature map and the sound features possibly related to the core heard from the dynamic voiceprint features, and couple and process this information on different aspects of the core status through a specific algorithm. For example, if it is found that the core vibration frequency has a slight change during a certain period and the corresponding sound also becomes slightly dull, after eigen-coupling analysis, the server concludes that the core is in a transient state with a possible slight abnormality during this period, which is part of the transient device status representation. Similarly, such analysis is also performed on other device status components such as the winding status, so as to obtain a complete transient device status representation. After obtaining the transient device status representation, the server then performs cross-modal coupling on the transient device status representation, the dynamic acoustic-vibration joint feature map, and the dynamic voiceprint features. Still taking the transformer as an example, the server will deeply fuse again the transient status information of various aspects such as the core and winding in the transient device status representation with the vibration and voiceprint information at different frequencies and amplitudes accurately presented in the dynamic acoustic-vibration joint feature map and the detailed sound features in the dynamic voiceprint features. Through calculation, these information from different modalities (sound and vibration) and different levels are comprehensively integrated to obtain a dynamic device status representation that can more accurately reflect the current actual operating status of the transformer. Finally, the server needs to calibrate the obtained dynamic device status representation to the time-frequency representation domain of the preset device status representation. For example, the time-frequency representation domain of the preset device status representation has its specific frequency range and time scale standard. The server will adjust and adapt the information in the dynamic device status representation according to these standards to make it meet the requirements of the preset time-frequency representation domain. After calibration, an intermediate device status representation is obtained, which is a status representation that can more accurately reflect the current operating status of the transformer after comprehensively considering various information and undergoing standardized calibration.

[0072] In an embodiment of the present invention, the time-frequency analysis of the to-be-monitored voiceprint segment to obtain voiceprint features at multiple resolution levels can be implemented through the following examples.

[0073] Perform multi-scale time-frequency analysis on the to-be-monitored voiceprint segment to obtain initial voiceprint features at at least one frequency band decomposition level;

[0074] Perform time-frequency calibration on the initial voiceprint features to obtain transient voiceprint features at multiple resolution levels;

[0075] According to a preset coupling topology, perform cross-level coupling on the transient voiceprint features to obtain voiceprint features at multiple resolution levels.

[0076] In an embodiment of the present invention, by way of example, in a large power plant, a server is responsible for monitoring the operating status of a generator set. For the voiceprint segments to be monitored of the generator set collected, the server will perform multi-scale time-frequency analysis using an advanced time-frequency analysis algorithm. For example, taking a large steam turbine generator set as an example, the sound emitted during its operation contains various frequency components and changes over time. Through multi-scale time-frequency analysis, the server dissects the voiceprint segments at different time scales and frequency scales, just like observing the details of the sound with magnifying glasses of different precisions. In this way, the server can obtain initial voiceprint features at at least one frequency band decomposition level, and each frequency band decomposition level corresponds to the sound characteristics of different frequency ranges. For example, the low frequency band may be related to the sound associated with the mechanical vibration of the unit, and the high frequency band may be related to the sound associated with the electromagnetic vibration of electrical components. After obtaining the initial voiceprint features, the server then performs time-frequency calibration on them. Still taking this steam turbine generator set as an example, since there may be some deviations in the time and frequency correspondence relationships of the initial voiceprint features at different frequency band decomposition levels, the server adjusts the time and frequency parameters of the initial voiceprint features at each level through a specific calibration algorithm to make them more regular and accurate in the time-frequency domain. After time-frequency calibration, the server obtains transient voiceprint features at multiple resolution levels, and these transient voiceprint features can more accurately reflect the transient characteristics of the sound of the generator set at different times and frequencies. Finally, the server performs cross-level coupling on the transient voiceprint features according to the preset coupling topology. Assume that the preset coupling topology stipulates the coupling method and weight between transient voiceprint features at different levels. The server performs deep fusion on the transient voiceprint features at different resolution levels according to this rule, just like piecing together different-level voice information puzzles in a specific way. Through this cross-level coupling, the server finally obtains voiceprint features at multiple resolution levels, and these voiceprint features comprehensively and accurately present various characteristics of the sound during the operation of the generator set, providing an important basis for subsequent monitoring of the equipment status.

[0077] In an embodiment of the present invention, the time-frequency calibration of the initial voiceprint features to obtain transient voiceprint features at multiple resolution levels can be implemented through the following example.

[0078] Based on the frequency band decomposition level, locate the target initial voiceprint feature in the initial voiceprint features, and perform time-frequency grid alignment on the target initial voiceprint feature to obtain the voiceprint features with alignment completed;

[0079] Perform time-frequency compounding on the voiceprint features with alignment completed and the installation topology features to obtain compound voiceprint features, and the compound voiceprint features include at least one voiceprint sub-feature;

[0080] Perform eigen-coupling analysis on the voiceprint sub-features to obtain voiceprint features with completed energy ratio, and perform dynamic reconstruction on the initial voiceprint features based on the voiceprint features with completed energy ratio to obtain transient voiceprint features at multiple resolution levels.

[0081] In an embodiment of the present invention, for example, in the scenario of monitoring power equipment in a substation, the server is responsible for processing the voiceprint data of the transformer. For the obtained initial voiceprint features, there are multiple frequency band decomposition levels. The server accurately locates the target initial voiceprint features based on these frequency band decomposition levels. For example, according to the operating characteristics of the transformer, the initial voiceprint features corresponding to the frequency range related to the core vibration are selected as the target initial voiceprint features. Then, a time-frequency grid alignment operation is performed on the target initial voiceprint features. It is like tidying up the uneven "data grids" to make them neat and tidy. By using a specific algorithm to adjust their positions and intervals in the time-frequency domain, the division of time and frequency becomes more standardized and accurate, thus obtaining the voiceprint features with completed alignment. Next, the server performs time-frequency compounding on the voiceprint features with completed alignment and the installation topology features. Taking the transformer as an example, the installation topology features include information such as the installation positions and connection relationships of various components of the transformer. Performing time-frequency compounding on the two is like combining the sound information with the physical structure information of the equipment. After this operation, composite voiceprint features are obtained, which contain at least one voiceprint sub-feature. These voiceprint sub-features can reflect the relationship between the sound characteristics during the operation of the transformer and the equipment structure from different angles. Finally, the server performs eigen-coupling analysis on the voiceprint sub-features. For each voiceprint sub-feature, analyze its inherent characteristics such as frequency and amplitude and the relationships between them. Through eigen-coupling analysis, voiceprint features with completed energy ratio are obtained, and then the initial voiceprint features are dynamically reconstructed based on this feature. It is like recombining the data according to a new "formula", thereby obtaining transient voiceprint features at multiple resolution levels, which can more accurately reflect the operating sound state of the transformer at different times.

[0082] In an embodiment of the present invention, the preset coupling topology includes a main coupling path and a secondary coupling path. The main coupling path preferentially processes the fundamental frequency and harmonic levels of mechanical vibration, and the secondary coupling path processes the high-frequency levels of electromagnetic noise and partial discharge. According to the preset coupling topology, cross-level coupling is performed on the transient voiceprint features to obtain voiceprint features at multiple resolution levels, which can be implemented through the following examples.

[0083] Based on the frequency band energy levels of the transient voiceprint features, perform an energy level sorting operation on the transient voiceprint features to obtain a first frequency band energy level sequence;

[0084] According to the first frequency band energy level sequence, perform cross-level coupling on the transient voiceprint features along the main coupling path to obtain at least one transient coupling mode;

[0085] Perform an energy level sorting operation on the transient coupling mode based on the frequency band energy level of the transient coupling mode to obtain a second frequency band energy level sequence;

[0086] According to the second frequency band energy level sequence, perform cross-level coupling on the transient coupling mode along the secondary coupling path to obtain at least one coupling mode, and use each coupling mode as the voiceprint feature of a resolution level.

[0087] In an embodiment of the present invention, exemplarily, in the power equipment monitoring work of a large substation, the server is responsible for processing relevant data of the transformer. For the obtained transient acoustic fingerprint features, the server will analyze the frequency band energy level conditions. For example, the sound emitted during the operation of the transformer has different energy manifestations corresponding to different frequency bands. The server performs an energy level sorting operation on these transient acoustic fingerprint features according to the energy level size of each frequency band of the transient acoustic fingerprint features through a specific algorithm. Just like arranging the sound frequency bands in order of energy level from high to low, the first frequency band energy level sequence is obtained. Among them, there may be situations where the energy levels of the frequency bands related to mechanical vibration in the low frequency band are relatively high, and the energy levels of the frequency bands related to electromagnetic noise in the high frequency band are relatively low, etc. The server then performs cross-level coupling on the transient acoustic fingerprint features according to the first frequency band energy level sequence along the main coupling path. Since the main coupling path preferentially processes the fundamental frequency and harmonic levels of mechanical vibration, the server will focus on the transient acoustic fingerprint features corresponding to the frequency bands related to mechanical vibration in the first frequency band energy level sequence. Taking the core vibration of the transformer as an example, the fundamental frequency and harmonics of the mechanical vibration generated by it are correspondingly reflected in the acoustic fingerprint features. The server performs deep fusion on these transient acoustic fingerprint features at different levels related to mechanical vibration through a specific coupling algorithm, just like splicing different levels but related mechanical vibration sound information. After such cross-level coupling operation, the server obtains at least one transient coupling mode, and these transient coupling modes more concentratedly reflect the acoustic fingerprint feature information in terms of mechanical vibration. After obtaining the transient coupling mode, the server performs an energy level sorting operation on the transient coupling mode based on its frequency band energy level. Similarly, through a specific algorithm, it rearranges according to the energy level size of each frequency band in the transient coupling mode to obtain the second frequency band energy level sequence. At this time, after being processed by the main coupling path, the energy levels of some frequency bands related to electromagnetic noise may have new relative positions in the transient coupling mode. Finally, the server performs cross-level coupling on the transient coupling mode according to the second frequency band energy level sequence along the secondary coupling path. Since the secondary coupling path processes the electromagnetic noise and high frequency levels of partial discharge, the server will re-fuse the frequency bands in the transient coupling mode related to electromagnetic noise and partial discharge according to the new energy level sequence. For example, for the high frequency electromagnetic noise generated by the possible partial discharge of the transformer, the server further integrates the relevant acoustic fingerprint features through the cross-level coupling operation of the secondary coupling path. After such processing, the server obtains at least one coupling mode, and uses each coupling mode as the acoustic fingerprint feature of a resolution level. These acoustic fingerprint features comprehensively and specifically reflect the sound characteristics of different aspects during the operation of the transformer, providing a strong basis for monitoring the equipment status.

[0088] In an embodiment of the present invention, the step of performing cross-level coupling on the transient acoustic fingerprint features according to the first frequency band energy level sequence along the main coupling path to obtain at least one transient coupling mode can be implemented through the following examples.

[0089] Align to the transient voiceprint feature with the lowest frequency band energy level among the transient voiceprint features to obtain a dynamic transient voiceprint feature, and use the dynamic transient voiceprint feature as the first transient coupling mode;

[0090] Based on the first frequency band energy level sequence, align to the next transient voiceprint feature of the dynamic transient voiceprint feature in the transient voiceprint features to obtain the first coupling mode to be coupled of the dynamic transient voiceprint feature;

[0091] Perform time-frequency compounding on the dynamic transient voiceprint feature and the first coupling mode to be coupled to obtain a second transient coupling mode, and use the first coupling mode to be coupled as the dynamic transient voiceprint feature;

[0092] Recursively trigger the process of aligning to the next transient voiceprint feature of the dynamic transient voiceprint feature in the transient voiceprint features based on the first frequency band energy level sequence, and terminate the recursive process until all transient voiceprint features complete cross-level coupling, to obtain at least one transient coupling mode, where the transient coupling mode includes the first transient coupling mode and the second transient coupling mode.

[0093] In an embodiment of the present invention, for example, when a substation monitors a power transformer, the server is responsible for processing relevant voiceprint data. For the obtained transient voiceprint features, the server operates according to their frequency band energy levels. The server accurately locates the transient voiceprint feature with the lowest frequency band energy level in the transient voiceprint features. For example, when the transformer is operating, the transient voiceprint features corresponding to some low-frequency and weak-energy sounds may be the ones with the lowest frequency band energy level. It is determined as the dynamic transient voiceprint feature and used as the first transient coupling mode. This first transient coupling mode is like a basic module for constructing subsequent couplings, initially capturing the sound features at a specific frequency band energy level. Then, based on the first frequency band energy level sequence, the server continues to locate in the transient voiceprint features to find the next transient voiceprint feature of the dynamic transient voiceprint feature, which is used as the first coupling mode to be coupled of the dynamic transient voiceprint feature. Taking the transformer as an example, if the first transient coupling mode mainly reflects the low-frequency sound features generated by the slight vibration of the iron core, then the first coupling mode to be coupled may correspond to the relevant sound features with slightly different frequencies or energy manifestations from this low-frequency sound, such as the sound features related to the harmonics of the iron core vibration. Then, the server performs time-frequency compounding on the dynamic transient voiceprint feature and the first coupling mode to be coupled. This is like fusing two somewhat related but different sound information according to the rules of time and frequency. After being processed by a specific algorithm, the second transient coupling mode is obtained. At this time, the second transient coupling mode combines the characteristics of the previous two and can more comprehensively reflect some of the sound conditions related to mechanical vibration. After that, the server recursively triggers the process of locating the next transient voiceprint feature of the dynamic transient voiceprint feature in the transient voiceprint features based on the first frequency band energy level sequence. Each time it is triggered, the above operations of finding the coupling mode to be coupled and performing time-frequency compounding are repeated. For example, as the recursion progresses, new coupling modes to be coupled are continuously found and compounded with the current dynamic transient voiceprint feature, gradually integrating more relevant sound features. Until all transient voiceprint features complete cross-level coupling, the recursive process terminates, and the server obtains at least one transient coupling mode. These transient coupling modes (including the first transient coupling mode and the second transient coupling mode, etc.) can more completely present the voiceprint feature information related to mechanical vibration in the main coupling path, providing a richer basis for further monitoring the transformer status.

[0094] In an embodiment of the present invention, the operation of performing time-frequency compounding on the dynamic transient voiceprint feature and the coupling mode to be coupled to obtain the second transient coupling mode can be executed through the following example.

[0095] Perform frequency band dimension reallocation on the dynamic transient voiceprint feature to obtain the reallocated voiceprint features with a preset number of dimensions;

[0096] Perform energy level gain compensation on the frequency band energy levels of the reconfigured voiceprint features to obtain voiceprint features after energy level gain compensation, where the frequency band energy levels of the voiceprint features after energy level gain compensation are consistent with the frequency band energy levels of the first mode to be coupled;

[0097] Perform frequency band dimension reconfiguration on the first mode to be coupled to obtain the reconfigured first mode to be coupled with the preset number of dimensions;

[0098] Perform cross-frequency band compounding on the voiceprint features after energy level gain compensation and the reconfigured first mode to be coupled to obtain a second transient coupling mode.

[0099] In an embodiment of the present invention, by way of example, in the scenario of power equipment monitoring in a large power plant, the server is responsible for processing the voiceprint data of numerous generator sets. Taking one large steam turbine generator set as an example, the server has obtained the relevant dynamic and transient voiceprint features. These dynamic and transient voiceprint features record the sound characteristics of the generator set at a certain stage during operation, but their frequency band dimensions may not meet the requirements of subsequent composite operations. Therefore, the server will perform frequency band dimension reconfiguration on the dynamic and transient voiceprint features. For example, the original dynamic and transient voiceprint features may be relatively coarsely divided in the frequency band, or the information in some key frequency bands is not prominent enough. The server uses specific algorithms and data processing technologies to re-adjust its frequency band dimensions according to preset rules. Suppose the original dynamic and transient voiceprint features have three sub-frequency bands in the low-frequency band and two sub-frequency bands in the high-frequency band. After reconfiguration, according to the preset number of dimensions, it is adjusted to four sub-frequency bands in the low-frequency band and three sub-frequency bands in the high-frequency band, thus obtaining the reconfigured voiceprint features with the preset number of dimensions. Such reconfiguration makes the voiceprint features more regular in the frequency band dimension, which is more conducive to subsequent composite operations with other features, so as to more accurately reflect the characteristics of the sound of the generator set at different frequency bands during operation. After completing the frequency band dimension reconfiguration, although the obtained reconfigured voiceprint features meet the requirements in the frequency band dimension, the frequency band energy levels may not be consistent with those of the first to-be-coupled mode to be compounded. Still taking this steam turbine generator set as an example, the server will perform energy level gain compensation on the frequency band energy levels of the reconfigured voiceprint features. For example, the energy level of the reconfigured voiceprint features in a certain specific frequency band is relatively low, while the energy level of the corresponding first to-be-coupled mode in this frequency band is high. In order to make the two match in the frequency band energy levels, the server will calculate the energy value to be compensated according to the difference in their energy levels through a specific algorithm, and then perform an energy addition operation on this frequency band of the reconfigured voiceprint features. After such energy level gain compensation, the voiceprint features with energy level gain compensation are obtained, and at this time their frequency band energy levels are consistent with those of the first to-be-coupled mode. This is like adjusting the corresponding parts of two "sound puzzles" to be spliced together to the same "brightness" (energy level) so that subsequent composite operations can be carried out more perfectly to accurately reflect the correlation of the sound characteristics of the generator set in this frequency band. At the same time, the server also needs to perform frequency band dimension reconfiguration on the first to-be-coupled mode. For a steam turbine generator set, the first to-be-coupled mode may record another part of the sound characteristics that are related to but different from the dynamic and transient voiceprint features. Its original frequency band dimension may also be unsuitable for composite operations. The server performs frequency band dimension reconfiguration on the first to-be-coupled mode through a specific algorithm according to the same preset number of dimensions as the frequency band dimension reconfiguration of the dynamic and transient voiceprint features.For example, the first mode to be coupled originally had two sub-bands in the mid-frequency band. After reconfiguration, it was adjusted to three sub-bands in the mid-frequency band according to the preset number of dimensions, thus obtaining the first mode to be coupled after reconfiguration with the preset number of dimensions. Such reconfiguration makes the first mode to be coupled also more matched with the voiceprint features after reconfiguration in terms of the frequency band dimension, laying a good foundation for the subsequent cross-frequency band compound operation. Finally, the server performs cross-frequency band compounding on the voiceprint features after energy level gain compensation and the first mode to be coupled after reconfiguration. Taking a steam turbine generator set as an example, after a series of previous operations, the voiceprint features after energy level gain compensation and the first mode to be coupled after reconfiguration have reached a relatively matched state in terms of the frequency band dimension and the frequency band energy level. The server deeply fuses the information of these two features in different frequency bands through a specific compounding algorithm. For example, in the low-frequency band, the low-frequency band information of the voiceprint features after energy level gain compensation is merged with the low-frequency band information of the first mode to be coupled after reconfiguration according to certain weights and rules; in the high-frequency band, the high-frequency band information of both is similarly fused. Through such a cross-frequency band compound operation, the two features that originally recorded different sound characteristics are integrated together to obtain the second transient coupling mode. This second transient coupling mode combines the advantages of both and can more comprehensively and accurately reflect the sound characteristics related to this stage during the operation of the steam turbine generator set, providing a richer and more accurate basis for further analyzing the operating state of the generator set. When monitoring a power transformer in a substation, the situation is similar. After the server obtains the dynamic transient voiceprint features and the first mode to be coupled of the transformer, it will also perform operations such as frequency band dimension reconfiguration, energy level gain compensation, and cross-frequency band compounding according to the above steps, and finally obtain the second transient coupling mode to better analyze the operating state of the transformer, such as judging whether there are situations that may affect the normal operation of the equipment, such as loose iron core and overheated winding. Through such a detailed processing process, the server can more effectively utilize the voiceprint data to monitor the operating state of power equipment, timely discover potential problems, and ensure the stable operation of the power system.

[0100] In the embodiment of the present invention, the step of performing cross-level coupling on the transient coupling mode according to the second frequency band energy level sequence along the secondary coupling path to obtain at least one coupling mode can be implemented through the following examples.

[0101] Among the transient coupling modes, the transient coupling mode with the highest frequency band energy level is located to obtain the dynamic coupling mode, and the features in different frequency band dimensions in the dynamic coupling mode are subjected to time-frequency compounding to obtain the first coupling mode;

[0102] Based on the second frequency band energy level sequence, in the transient coupling mode, the next transient coupling mode of the dynamic coupling mode is located to obtain the second mode to be coupled of the dynamic coupling mode;

[0103] Perform time-frequency compounding on the dynamic coupling mode and the second mode to be coupled to obtain a second coupling mode, and use the second mode to be coupled as the dynamic coupling mode;

[0104] Recursively trigger the process of the next transient coupling mode in the transient coupling mode aligned with the dynamic coupling mode based on the second frequency band energy level sequence until the recursive process terminates when all transient coupling modes complete cross-level coupling, obtaining at least one coupling mode, where the coupling mode includes the first coupling mode and the second coupling mode.

[0105] In an embodiment of the present invention, exemplarily, when the server monitors a power transformer in a substation, relevant transient coupling mode data has been obtained. Based on the second frequency band energy level sequence, the server accurately locates the transient coupling mode with the highest frequency band energy level in the transient coupling mode. For example, when the transformer is operating, the transient coupling mode related to some high-frequency and high-energy electromagnetic noises may be the one with the highest frequency band energy level. It is determined as the dynamic coupling mode. Then, the server performs time-frequency compounding on the features of different frequency band dimensions in the dynamic coupling mode. Just like fusing the sound features regarding electromagnetic noises and other aspects under different frequency band dimensions according to the time and frequency rules, after being processed by a specific algorithm, the first coupling mode is obtained. This first coupling mode preliminarily integrates the key information of each frequency band dimension in the dynamic coupling mode and can more centrally reflect the partial sound situation related to high-frequency electromagnetism. Next, based on the second frequency band energy level sequence, the server locates the next transient coupling mode of the dynamic coupling mode in the transient coupling mode and uses it as the second coupling mode to be coupled of the dynamic coupling mode. For example, if the first coupling mode mainly reflects the high-frequency electromagnetic noise characteristics generated at the initial stage of partial discharge, then the second coupling mode to be coupled may correspond to the sound characteristics related to the subsequent changes of the high-frequency electromagnetic noise. Subsequently, the server performs time-frequency compounding on the dynamic coupling mode and the second coupling mode to be coupled. This is like fusing the existing comprehensive features with the new relevant features again according to the time and frequency rules to obtain the second coupling mode. At this time, the second coupling mode further synthesizes the characteristics of both and can more comprehensively reflect the sound situation related to electromagnetic noises and other aspects. After that, the server will recursively trigger the process of locating the next transient coupling mode of the dynamic coupling mode in the transient coupling mode based on the second frequency band energy level sequence. Each time it is triggered, the operations of finding the second coupling mode to be coupled and performing time-frequency compounding will be repeated. For example, as the recursion progresses, new second coupling modes to be coupled will be continuously found and compounded with the current dynamic coupling mode, gradually integrating more relevant sound features. Until all transient coupling modes complete cross-level coupling, the recursive process terminates, and the server obtains at least one coupling mode. These coupling modes (including the first coupling mode and the second coupling mode, etc.) can more completely present the voiceprint feature information related to electromagnetic noises and other aspects of the sub-coupling path, providing a richer basis for further monitoring the transformer status.

[0106] In addition, in a large-scale substation, the server is responsible for comprehensively monitoring numerous power equipment. Taking a large transformer as an example, in addition to analyzing the coupling coefficient of its acoustic fingerprint characteristics and vibration characteristics, the server also real-time monitors the current operating environment parameters of the transformer. The server accurately obtains information such as its temperature, humidity, and electromagnetic interference intensity through various sensors installed near the transformer. For example, in the hot summer, the temperature around the transformer may rise to 40°C or even higher, the humidity will also fluctuate due to weather changes, and at the same time, the operation of nearby electrical equipment may generate different degrees of electromagnetic interference. According to the preset adjustment rules for environmental parameters and coupling coefficients, when the server monitors an increase in temperature, according to the rules, the correlation between the acoustic fingerprint and vibration characteristics related to temperature changes in certain time-frequency units may be affected. For example, high temperature may cause some components inside the transformer to expand and contract thermally, thereby changing its vibration characteristics and affecting the coupling relationship with the acoustic fingerprint characteristics. At this time, the server will dynamically adjust the coupling coefficients in each time-frequency unit based on the changes in parameters such as the real-time monitored temperature, humidity, and electromagnetic interference intensity. It may reduce the coupling coefficients of those time-frequency units whose correlation weakens due to temperature changes to more accurately reflect the true operating state of the transformer in the current environment.

[0107] Still taking this large transformer as an example, after obtaining the reference acoustic-vibration combined feature maps at the same resolution level, the server needs to perform time-frequency compounding on them to obtain the complete acoustic-vibration combined feature maps. The server takes these reference acoustic-vibration combined feature maps as input data and inputs them into a pre-trained deep learning model. This deep learning model is carefully trained based on a large amount of labeled acoustic-vibration combined feature map data of power equipment in normal and abnormal states. When the input data enters the model, it first passes through the convolutional layer. The convolutional layer is like a delicate filter that slides a specific convolutional kernel over the data to extract features at different positions and scales. For example, it can identify the local features related to the vibration frequency of the transformer core in the reference acoustic-vibration combined feature map. Then, the data enters the pooling layer. The pooling layer compresses and simplifies the features extracted by the convolutional layer, removes some redundant information, and retains the key features. For example, it merges some adjacent and similar features to reduce the data volume without losing important information. Finally, it passes through the fully connected layer structure. The fully connected layer comprehensively fuses the features processed previously, recombines them according to the correlation relationships between the features, and finally outputs the acoustic-vibration combined feature maps at each resolution level. These maps can more comprehensively and accurately reflect the comprehensive acoustic fingerprint and vibration information of the transformer at different resolution levels.

[0108] After the server performs cross-modal coupling on the transient device state characterization, dynamic acoustic vibration joint feature map, and dynamic voiceprint features of the transformer to obtain the dynamic device state characterization, it is necessary to verify its reliability. The server will carefully extract the confirmed and accurate device state characterization data under similar operating conditions from the historical power equipment operation data as the verification sample set. For example, find the accurately recorded device state characterization data when the transformer was operating under similar conditions such as temperature and load in the past. Then, using the similarity measurement algorithm, calculate the similarity between the dynamic device state characterization and each sample in the verification sample set. Assume that the preset reliability threshold is 80%. If the calculated similarity is lower than 80%, it indicates that the currently obtained dynamic device state characterization may not be accurate enough. At this time, the server will re-execute the process of detecting at least one power equipment state to be detected in the voiceprint segment to be monitored based on the acoustic vibration joint feature map and voiceprint features. After a series of analyses and processes again, continuously adjust and improve until the obtained dynamic device state characterization passes the reliability verification to ensure the accuracy of the judgment of the transformer operating state.

[0109] During the process of processing the transformer-related data, for each transient coupling mode, the server will perform further operations according to its corresponding current frequency band energy level. For example, the current frequency band energy level corresponding to a certain transient coupling mode is at a medium level. According to the preset mapping relationship between the energy level and the enhancement coefficient, the server determines that its corresponding feature enhancement coefficient is 1.2 (assumed). Then, use this feature enhancement coefficient to weight the features of the transient coupling mode in the frequency band dimension. It is like "amplifying" the different features of this transient coupling mode in the frequency band dimension according to a ratio of 1.2, so that these features can more prominently show their importance in subsequent analyses and processes and more accurately reflect the operating state of the transformer in aspects related to this frequency band energy level.

[0110] An embodiment of the present invention provides a power equipment monitoring system based on multi-modal information collaboration, including:

[0111] An initial unit for obtaining the voiceprint to be monitored and the vibration information to be analyzed corresponding to the voiceprint to be monitored, where the voiceprint to be monitored includes multiple voiceprint segments to be monitored;

[0112] An analysis unit for performing time-frequency analysis on each voiceprint segment to be monitored in the multiple voiceprint segments to be monitored to obtain voiceprint features at multiple resolution levels, and performing time-frequency analysis on the vibration information to be analyzed to obtain vibration features;

[0113] A calculation unit for calculating the coupling coefficient between the voiceprint features and the vibration features in each time-frequency unit, where the coupling coefficient reflects the correlation strength between the voiceprint features and the vibration features in the same time-frequency unit;

[0114] A composite unit, configured to perform time-frequency composition on the voiceprint feature and the vibration feature based on the coupling coefficient to obtain a voice-vibration combined feature map;

[0115] A detection unit, configured to detect at least one power equipment state to be detected in the voiceprint segment to be monitored according to the voice-vibration combined feature map and the voiceprint feature.

[0116] Wherein, the vibration feature includes at least one vibration sub-mode, and calculating the coupling coefficient between the voiceprint feature and the vibration feature in each time-frequency unit includes:

[0117] Identifying the number of frequency band dimensions in the voiceprint feature to obtain the voiceprint frequency band number, and identifying the number of frequency band dimensions in the vibration feature to obtain the vibration frequency band number;

[0118] Based on the voiceprint frequency band number and the vibration frequency band number, determining the target frequency band dimension number of the time-frequency representation domain corresponding to each resolution level, and performing time-domain superposition on the vibration sub-modes to obtain a superimposed vibration feature;

[0119] Optimizing the dimension numbers of the voiceprint feature and the superimposed vibration feature to the target frequency band dimension number respectively to obtain the aligned voiceprint feature corresponding to each resolution level and the aligned vibration feature corresponding to the aligned voiceprint feature;

[0120] Extracting the voiceprint sub-features corresponding to each time-frequency unit from the aligned voiceprint feature;

[0121] Extracting the features under each frequency band dimension from the voiceprint sub-features to obtain voiceprint frequency band features;

[0122] Aligning the features in the aligned vibration feature to the frequency band dimension corresponding to the voiceprint frequency band feature to obtain vibration frequency band features;

[0123] Coupling the voiceprint frequency band features and the vibration frequency band features to obtain the coupling channel responses corresponding to each frequency band dimension;

[0124] Integrating the coupling channel responses in each time-frequency unit to obtain the joint coupling responses corresponding to each time-frequency unit;

[0125] Coupling the voiceprint sub-features and the aligned vibration features to obtain a reference coupling response, and calculating the energy ratio between the reference coupling response and the joint coupling response to obtain the coupling coefficient corresponding to each time-frequency unit.

[0126] Wherein, performing time-frequency composition on the voiceprint feature and the vibration feature based on the coupling coefficient to obtain a voice-vibration combined feature map includes:

[0127] Based on the coupling coefficient, perform energy ratio matching on the voiceprint sub-features to obtain a reference acoustic-vibration joint feature map corresponding to each time-frequency unit;

[0128] Perform time-frequency compounding on the reference acoustic-vibration joint feature maps at the same resolution level to obtain acoustic-vibration joint feature maps at each resolution level.

[0129] Among them, according to the acoustic-vibration joint feature map and the voiceprint feature, detecting at least one power equipment state to be detected in the voiceprint segment to be monitored includes:

[0130] Based on the resolution level of the voiceprint feature, perform hierarchical optimization on the voiceprint feature, and align the voiceprint feature at the target resolution level in the voiceprint feature based on the optimized hierarchical sequence to obtain a dynamic voiceprint feature;

[0131] Align to the acoustic-vibration joint feature map corresponding to the target resolution level in the acoustic-vibration joint feature map to obtain a dynamic acoustic-vibration joint feature map;

[0132] Perform intrinsic coupling analysis on the device state component to obtain a transient device state representation;

[0133] Perform cross-modal coupling on the transient device state representation, the dynamic acoustic-vibration joint feature map, and the dynamic voiceprint feature to obtain a dynamic device state representation;

[0134] Calibrate the dynamic device state representation to the time-frequency representation domain of the preset device state representation to obtain an intermediate device state representation, and use the intermediate device state representation as the preset device state representation; the preset device state representation includes at least one device state component,

[0135] Recursively trigger the process from performing intrinsic coupling analysis on the device state component to obtain a transient device state representation to calibrating the dynamic device state representation to the time-frequency representation domain of the preset device state representation to obtain an intermediate device state representation, until the recursive process terminates when a preset convergence criterion is reached, to obtain an incrementally corrected device state representation, the preset convergence criterion includes a device state representation difference rate threshold and a phase coherence coefficient threshold, and use the incrementally corrected device state representation as the preset device state representation;

[0136] Recursively trigger the process of aligning the voiceprint feature at the target resolution level in the voiceprint feature based on the optimized hierarchical sequence until the recursive process terminates when each voiceprint feature is the dynamic voiceprint feature, to obtain a device state representation, and based on the device state representation, determine at least one power equipment state to be detected in the voiceprint segment to be monitored.

[0137] Among them, time-frequency analysis is performed on the to-be-monitored voiceprint segment to obtain voiceprint features at multiple resolution levels, including:

[0138] Perform multi-scale time-frequency analysis on the to-be-monitored voiceprint segment to obtain initial voiceprint features at at least one frequency band decomposition level;

[0139] Based on the frequency band decomposition level, align to the target initial voiceprint feature in the initial voiceprint features, and perform time-frequency grid alignment on the target initial voiceprint feature to obtain the voiceprint features with alignment completed;

[0140] Perform time-frequency compounding on the voiceprint features with alignment completed and the installation topology features to obtain compound voiceprint features, and the compound voiceprint features include at least one voiceprint sub-feature;

[0141] Perform eigen-coupling analysis on the voiceprint sub-feature to obtain voiceprint features with energy ratio completed, and perform dynamic reconstruction on the initial voiceprint features based on the voiceprint features with energy ratio completed to obtain transient voiceprint features at multiple resolution levels;

[0142] Based on the frequency band energy levels of the transient voiceprint features, perform an energy level sorting operation on the transient voiceprint features to obtain a first frequency band energy level sequence;

[0143] Align to the transient voiceprint feature with the lowest frequency band energy level in the transient voiceprint features to obtain a dynamic transient voiceprint feature, and use the dynamic transient voiceprint feature as the first transient coupling mode;

[0144] Based on the first frequency band energy level sequence, align to the next transient voiceprint feature of the dynamic transient voiceprint feature in the transient voiceprint features to obtain the first to-be-coupled mode of the dynamic transient voiceprint feature;

[0145] Perform frequency band dimension reallocation on the dynamic transient voiceprint feature to obtain reallocated voiceprint features with a preset dimension number;

[0146] Perform energy level gain compensation on the frequency band energy levels of the reallocated voiceprint features to obtain voiceprint features after energy level gain compensation, and the frequency band energy levels of the voiceprint features after energy level gain compensation are consistent with the frequency band energy levels of the first to-be-coupled mode;

[0147] Perform frequency band dimension reallocation on the first to-be-coupled mode to obtain the reallocated first to-be-coupled mode with the preset dimension number;

[0148] Perform cross-frequency band compounding on the voiceprint features after energy level gain compensation and the reallocated first to-be-coupled mode to obtain a second transient coupling mode;

[0149] Use the first to-be-coupled mode as the dynamic transient voiceprint feature;

[0150] Recursively trigger the process of aligning the next transient voiceprint feature of the dynamic transient voiceprint feature in the transient voiceprint feature based on the first frequency band energy level sequence, and terminate the recursive process until all transient voiceprint features complete cross-level coupling, obtaining at least one transient coupling mode, where the transient coupling mode includes the first transient coupling mode and the second transient coupling mode;

[0151] Based on the frequency band energy level of the transient coupling mode, perform an energy level sorting operation on the transient coupling mode to obtain a second frequency band energy level sequence;

[0152] In the transient coupling mode, align to the transient coupling mode with the highest frequency band energy level to obtain a dynamic coupling mode, and perform time-frequency compounding on the features of different frequency band dimensions in the dynamic coupling mode to obtain a first coupling mode;

[0153] Based on the second frequency band energy level sequence, in the transient coupling mode, align to the next transient coupling mode of the dynamic coupling mode to obtain the second coupling mode to be coupled of the dynamic coupling mode;

[0154] Perform time-frequency compounding on the dynamic coupling mode and the second coupling mode to be coupled to obtain a second coupling mode, and use the second coupling mode to be coupled as the dynamic coupling mode;

[0155] Recursively trigger the process of aligning to the next transient coupling mode of the dynamic coupling mode in the transient coupling mode based on the second frequency band energy level sequence, and terminate the recursive process until all transient coupling modes complete cross-level coupling, obtaining at least one coupling mode, where the coupling mode includes the first coupling mode and the second coupling mode;

[0156] Use each coupling mode as the voiceprint feature of a resolution level.

[0157] Wherein, after calculating the coupling coefficient between the voiceprint feature and the vibration feature in each time-frequency unit, it further includes:

[0158] Real-time monitor the current operating environment parameters of the power equipment, where the current operating environment parameters include temperature, humidity, and electromagnetic interference intensity;

[0159] According to the preset adjustment rule of environmental parameters and coupling coefficient, dynamically adjust the coupling coefficient in each time-frequency unit based on the change of the currently monitored current operating environment parameters.

[0160] Wherein, performing time-frequency compounding on the reference acoustic-vibration joint feature maps of the same resolution level to obtain the acoustic-vibration joint feature maps of each resolution level includes:

[0161] Take the reference acoustic-vibration joint feature map at the same resolution level as input data and input it into a pre-trained deep learning model, where the deep learning model is trained based on the acoustic-vibration joint feature map data of power equipment in normal and abnormal states with annotations;

[0162] Extract and fuse features from the input data through the convolutional layer, pooling layer, and fully connected layer structures of the deep learning model, and output the acoustic-vibration joint feature maps at each resolution level.

[0163] Among them, after cross-modal coupling of the transient device state representation, the dynamic acoustic-vibration joint feature map, and the dynamic voiceprint feature to obtain the dynamic device state representation, it further includes:

[0164] Perform reliability verification on the dynamic device state representation, and extract the confirmed and accurate device state representation data under operating conditions similar to the current monitoring period from the historical power equipment operation data as the verification sample set;

[0165] Use a similarity measurement algorithm to calculate the similarity between the dynamic device state representation and each sample in the verification sample set;

[0166] If the similarity is lower than a preset reliability threshold, re-execute the process of detecting at least one power equipment state to be detected in the voiceprint segment to be monitored based on the acoustic-vibration joint feature map and the voiceprint feature until the obtained dynamic device state representation passes the reliability verification.

[0167] Among them, based on the frequency band energy level of the transient coupling mode, performing an energy level sorting operation on the transient coupling mode to obtain a second frequency band energy level sequence further includes:

[0168] For each transient coupling mode, determine the corresponding feature enhancement coefficient according to the current frequency band energy level corresponding to each according to the preset mapping relationship between the energy level and the enhancement coefficient;

[0169] Use the feature enhancement coefficient to perform weighted processing on the features of the transient coupling mode in the frequency band dimension.

[0170] An embodiment of the present invention provides a computer device 100. The computer device 100 includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the foregoing power equipment monitoring method based on multi-modal information collaboration. As Figure 2 shown, Figure 2It is a block diagram of a computer device 100 provided by an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To achieve data transmission or interaction, the elements of the memory 111, the processor 112, and the communication unit 113 are electrically connected to each other directly or indirectly. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.

[0171] An embodiment of the present invention provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, in this storage space, there are also stored one or more instructions suitable for being loaded and executed by the processor. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The one or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the steps of the method in the above embodiment.

[0172] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. The embodiments are chosen and described in order to best illustrate the principles of the disclosure and its practical applications, thereby enabling those skilled in the art to best utilize the disclosure and to utilize various embodiments with different modifications to suit the particular applications contemplated.

Claims

1. A power equipment monitoring method based on multi-modal information collaboration, characterized in that Including: Obtaining the voiceprint to be monitored and the vibration information to be analyzed corresponding to the voiceprint to be monitored, where the voiceprint to be monitored includes a plurality of voiceprint segments to be monitored; Performing time-frequency analysis on each voiceprint segment in the plurality of voiceprint segments to be monitored to obtain voiceprint features at multiple resolution levels, and performing time-frequency analysis on the vibration information to be analyzed to obtain vibration features; Calculating the coupling coefficient between the voiceprint features and the vibration features in each time-frequency unit, where the coupling coefficient reflects the correlation strength between the voiceprint features and the vibration features in the same time-frequency unit; Based on the coupling coefficient, performing time-frequency compounding on the voiceprint features and the vibration features to obtain a voice-vibration combined feature map; Detecting at least one power equipment state to be detected in the voiceprint segment to be monitored according to the voice-vibration combined feature map and the voiceprint features; 2. The method according to claim 1, characterized in that, The vibration features include at least one vibration sub-mode, and the calculating the coupling coefficient between the voiceprint features and the vibration features in each time-frequency unit includes: Identifying the number of frequency band dimensions in the voiceprint features to obtain the voiceprint frequency band number, and identifying the number of frequency band dimensions in the vibration features to obtain the vibration frequency band number; Based on the voiceprint frequency band number and the vibration frequency band number, determining the target frequency band dimension number of the time-frequency representation domain corresponding to each resolution level, and performing time-domain superposition on the vibration sub-modes to obtain the superimposed vibration features; Optimizing the dimension numbers of the voiceprint features and the superimposed vibration features to the target frequency band dimension number respectively to obtain the aligned voiceprint features corresponding to each resolution level and the aligned vibration features corresponding to the aligned voiceprint features; Extracting the voiceprint sub-features corresponding to each time-frequency unit from the aligned voiceprint features; Extracting the features under each frequency band dimension from the voiceprint sub-features to obtain the voiceprint frequency band features; Positioning the features under the frequency band dimension corresponding to the voiceprint frequency band features in the aligned vibration features to obtain the vibration frequency band features; Coupling the voiceprint frequency band features and the vibration frequency band features to obtain the coupling channel responses corresponding to each frequency band dimension; Integrating the coupling channel responses in each time-frequency unit to obtain the joint coupling responses corresponding to each time-frequency unit; Coupling the voiceprint sub-features and the aligned vibration features to obtain a reference coupling response, and calculating the energy ratio between the reference coupling response and the joint coupling response to obtain the coupling coefficient corresponding to each time-frequency unit.

3. The method according to claim 2, wherein The performing time-frequency compounding on the voiceprint features and the vibration features based on the coupling coefficient to obtain a voice-vibration combined feature map includes: Based on the coupling coefficient, performing energy ratio matching on the voiceprint sub-features to obtain the reference voice-vibration combined feature map corresponding to each time-frequency unit; Performing time-frequency compounding on the reference voice-vibration combined feature maps at the same resolution level to obtain the voice-vibration combined feature maps at each resolution level.

4. The method according to claim 1, characterized in that, The detecting at least one power equipment state to be detected in the voiceprint segment to be monitored according to the voice-vibration combined feature map and the voiceprint features includes: Based on the resolution levels of the voiceprint features, perform hierarchical optimization on the voiceprint features, and align the voiceprint features at the target resolution level in the voiceprint features based on the optimized hierarchical sequence to obtain dynamic voiceprint features; Align to the corresponding acoustic-vibration joint feature map at the target resolution level in the acoustic-vibration joint feature map to obtain a dynamic acoustic-vibration joint feature map; Perform eigen-coupling analysis on the device state components to obtain a transient device state representation; Perform cross-modal coupling on the transient device state representation, the dynamic acoustic-vibration joint feature map, and the dynamic voiceprint features to obtain a dynamic device state representation; Calibrate the dynamic device state representation to the time-frequency representation domain of the preset device state representation to obtain an intermediate device state representation, and use the intermediate device state representation as the preset device state representation; the preset device state representation includes at least one device state component, Recursively trigger the process from performing eigen-coupling analysis on the device state components to obtain a transient device state representation to calibrating the dynamic device state representation to the time-frequency representation domain of the preset device state representation to obtain an intermediate device state representation, and terminate the recursive process until the preset convergence criterion is reached. The preset convergence criterion includes a device state representation difference rate threshold and a phase coherence coefficient threshold, and use the incrementally corrected device state representation as the preset device state representation; Recursively trigger the process of aligning the voiceprint features at the target resolution level in the voiceprint features based on the optimized hierarchical sequence, and terminate the recursive process until all voiceprint features are the dynamic voiceprint features to obtain a device state representation, and determine at least one power device state to be detected in the voiceprint segment to be monitored based on the device state representation.

5. The method according to claim 1, characterized in that, The time-frequency analysis of the voiceprint segment to be monitored to obtain voiceprint features at multiple resolution levels includes: Perform multi-scale time-frequency analysis on the voiceprint segment to be monitored to obtain initial voiceprint features at at least one frequency band decomposition level; Based on the frequency band decomposition level, align to the target initial voiceprint feature in the initial voiceprint features, and perform time-frequency grid alignment on the target initial voiceprint feature to obtain the aligned voiceprint features; Perform time-frequency composition on the aligned voiceprint features and the installation topology features to obtain composite voiceprint features, and the composite voiceprint features include at least one voiceprint sub-feature; Perform eigen-coupling analysis on the voiceprint sub-features to obtain the voiceprint features with energy ratio completed, and perform dynamic reconstruction on the initial voiceprint features based on the voiceprint features with energy ratio completed to obtain transient voiceprint features at multiple resolution levels; Based on the frequency band energy levels of the transient voiceprint features, perform an energy level sorting operation on the transient voiceprint features to obtain a first frequency band energy level sequence; Align to the transient voiceprint feature with the lowest frequency band energy level in the transient voiceprint features to obtain a dynamic transient voiceprint feature, and use the dynamic transient voiceprint feature as the first transient coupling mode; Based on the first frequency band energy level sequence, the next transient voiceprint feature corresponding to the dynamic transient voiceprint feature is located in the transient voiceprint feature to obtain the first coupling mode to be coupled of the dynamic transient voiceprint feature; Perform frequency band dimension reassignment on the dynamic transient voiceprint feature to obtain the reassigned voiceprint feature with a preset number of dimensions; Perform energy level gain compensation on the frequency band energy level of the reassigned voiceprint feature to obtain the voiceprint feature after energy level gain compensation, and the frequency band energy level of the voiceprint feature after energy level gain compensation is consistent with the frequency band energy level of the first coupling mode to be coupled; Perform frequency band dimension reassignment on the first coupling mode to be coupled to obtain the reassigned first coupling mode to be coupled with the preset number of dimensions; Perform cross-frequency band compounding on the voiceprint feature after energy level gain compensation and the reassigned first coupling mode to be coupled to obtain the second transient coupling mode; Use the first coupling mode to be coupled as the dynamic transient voiceprint feature; Recursively trigger the process of locating the next transient voiceprint feature corresponding to the dynamic transient voiceprint feature in the transient voiceprint feature based on the first frequency band energy level sequence, and terminate the recursive process until all transient voiceprint features complete cross-level coupling, to obtain at least one transient coupling mode, where the transient coupling mode includes the first transient coupling mode and the second transient coupling mode; Based on the frequency band energy level of the transient coupling mode, perform an energy level sorting operation on the transient coupling mode to obtain the second frequency band energy level sequence; Locate the transient coupling mode with the highest frequency band energy level in the transient coupling mode to obtain the dynamic coupling mode, and perform time-frequency compounding on the features of different frequency band dimensions in the dynamic coupling mode to obtain the first coupling mode; Based on the second frequency band energy level sequence, locate the next transient coupling mode corresponding to the dynamic coupling mode in the transient coupling mode to obtain the second coupling mode to be coupled of the dynamic coupling mode; Perform time-frequency compounding on the dynamic coupling mode and the second coupling mode to be coupled to obtain the second coupling mode, and use the second coupling mode to be coupled as the dynamic coupling mode; Recursively trigger the process of locating the next transient coupling mode corresponding to the dynamic coupling mode in the transient coupling mode based on the second frequency band energy level sequence, and terminate the recursive process until all transient coupling modes complete cross-level coupling, to obtain at least one coupling mode, where the coupling mode includes the first coupling mode and the second coupling mode; Use each coupling mode as the voiceprint feature of a resolution level.

6. The method according to claim 1, characterized in that, After calculating the coupling coefficient between the voiceprint feature and the vibration feature in each time-frequency unit, it further includes: Real-time monitor the current operating environment parameters of the power equipment, where the current operating environment parameters include temperature, humidity, and electromagnetic interference intensity; According to the preset adjustment rule of environmental parameters and coupling coefficients, dynamically adjust the coupling coefficients in each time-frequency unit based on the change situation of the currently monitored current operating environment parameters.

7. The method according to claim 3, wherein The time-frequency compounding of the reference acoustic-vibration joint feature maps of the same resolution level to obtain the acoustic-vibration joint feature maps of each resolution level includes: Taking the reference acoustic-vibration joint feature map of the same resolution level as input data, input it into a pre-trained deep learning model, and the deep learning model is trained based on the acoustic-vibration joint feature map data of power equipment in normal and abnormal states with annotations; Through the convolutional layer, pooling layer and fully connected layer structure of the deep learning model, feature extraction and fusion processing are performed on the input data, and the acoustic-vibration joint feature maps of each resolution level are output.

8. The method according to claim 4, characterized in that After performing cross-modal coupling on the transient device state representation, the dynamic acoustic-vibration joint feature map and the dynamic voiceprint feature to obtain the dynamic device state representation, it further includes: Performing reliability verification on the dynamic device state representation, and extracting the confirmed accurate device state representation data under the operating conditions similar to the current monitoring period from the historical power equipment operation data as the verification sample set; Using a similarity metric algorithm to calculate the similarity between the dynamic device state representation and each sample in the verification sample set; If the similarity is lower than the preset reliability threshold, re-execute the process of detecting at least one power equipment state to be detected in the to-be-monitored voiceprint segment based on the acoustic-vibration joint feature map and the voiceprint feature until the obtained dynamic device state representation passes the reliability verification.

9. The method according to claim 5, wherein The performing an energy level sorting operation on the transient coupling mode based on the frequency band energy level of the transient coupling mode to obtain a second frequency band energy level sequence further includes: For each transient coupling mode, according to the corresponding current frequency band energy level, determine the corresponding feature enhancement coefficient according to the preset mapping relationship between the energy level and the enhancement coefficient; Using the feature enhancement coefficient to perform weighted processing on the features of the transient coupling mode in the frequency band dimension.

10. A power equipment monitoring system based on multi-modal information collaboration, characterized in that, It includes: An initial unit for obtaining a to-be-monitored voiceprint and the to-be-analyzed vibration information corresponding to the to-be-monitored voiceprint, where the to-be-monitored voiceprint includes a plurality of to-be-monitored voiceprint segments; An analysis unit for performing time-frequency analysis on each to-be-monitored voiceprint segment in the plurality of to-be-monitored voiceprint segments to obtain voiceprint features of multiple resolution levels, and performing time-frequency analysis on the to-be-analyzed vibration information to obtain vibration features; A calculation unit for calculating the coupling coefficient between the voiceprint feature and the vibration feature in each time-frequency unit, where the coupling coefficient reflects the correlation strength between the voiceprint feature and the vibration feature in the same time-frequency unit; A composite unit for performing time-frequency composition of the voiceprint feature and the vibration feature based on the coupling coefficient to obtain an acoustic-vibration joint feature map; A detection unit for detecting at least one power equipment state to be detected in the to-be-monitored voiceprint segment according to the acoustic-vibration joint feature map and the voiceprint feature.

11. The system according to claim 10, wherein The vibration feature includes at least one vibration sub-mode, and the calculating the coupling coefficient between the voiceprint feature and the vibration feature in each time-frequency unit includes: Identifying the number of frequency band dimensions in the voiceprint feature to obtain the voiceprint frequency band number, and identifying the number of frequency band dimensions in the vibration feature to obtain the vibration frequency band number; Based on the number of voiceprint frequency bands and the number of vibration frequency bands, determine the target frequency band dimension number of the time-frequency representation domain corresponding to each resolution level, and perform time-domain superposition on the vibration sub-modalities to obtain the superimposed vibration features; Optimize the dimension numbers of the voiceprint features and the superimposed vibration features to the target frequency band dimension number respectively to obtain the aligned voiceprint features at each resolution level and the aligned vibration features corresponding to the aligned voiceprint features; Extract the voiceprint sub-features corresponding to each time-frequency unit from the aligned voiceprint features; Extract the features under each frequency band dimension from the voiceprint sub-features to obtain the voiceprint frequency band features; Locate the features under the frequency band dimension corresponding to the voiceprint frequency band features in the aligned vibration features to obtain the vibration frequency band features; Couple the voiceprint frequency band features and the vibration frequency band features to obtain the coupled channel responses corresponding to each frequency band dimension; Integrate the coupled channel responses within each time-frequency unit to obtain the joint coupled responses corresponding to each time-frequency unit; Couple the voiceprint sub-features and the aligned vibration features to obtain the reference coupled response, and calculate the energy ratio between the reference coupled response and the joint coupled response to obtain the coupling coefficients corresponding to each time-frequency unit.

12. The system according to claim 11, wherein Based on the coupling coefficients, perform time-frequency compounding on the voiceprint features and the vibration features to obtain the acoustic-vibration joint feature map, including: Based on the coupling coefficients, perform energy ratio allocation on the voiceprint sub-features to obtain the reference acoustic-vibration joint feature maps corresponding to each time-frequency unit; Perform time-frequency compounding on the reference acoustic-vibration joint feature maps at the same resolution level to obtain the acoustic-vibration joint feature maps at each resolution level.

13. The system according to claim 10, characterized in that, According to the acoustic-vibration joint feature map and the voiceprint features, detect at least one power equipment state to be detected in the voiceprint segment to be monitored, including: Based on the resolution level of the voiceprint features, perform hierarchical optimization on the voiceprint features, and locate the voiceprint features at the target resolution level in the voiceprint features based on the optimized hierarchical sequence to obtain the dynamic voiceprint features; Locate the acoustic-vibration joint feature map corresponding to the target resolution level in the acoustic-vibration joint feature map to obtain the dynamic acoustic-vibration joint feature map; Perform intrinsic coupling analysis on the equipment state components to obtain the transient equipment state representation; Perform cross-modal coupling on the transient equipment state representation, the dynamic acoustic-vibration joint feature map, and the dynamic voiceprint features to obtain the dynamic equipment state representation; Calibrate the dynamic equipment state representation to the time-frequency representation domain of the preset equipment state representation to obtain the intermediate equipment state representation, and use the intermediate equipment state representation as the preset equipment state representation; the preset equipment state representation includes at least one equipment state component, Recursively trigger the process of performing eigen-coupling analysis on the device state components to obtain the transient device state representation, calibrate the dynamic device state representation to the time-frequency representation domain of the preset device state representation, and obtain the intermediate device state representation until the recursive process terminates when the preset convergence criterion is met, obtaining the device state representation after incremental correction. The preset convergence criterion includes the device state representation difference rate threshold and the phase coherence coefficient threshold, and use the device state representation after incremental correction as the preset device state representation; Recursively trigger the process of aligning the voiceprint features to the voiceprint features at the target resolution level in the voiceprint features based on the preferred hierarchical sequence until the recursive process terminates when each voiceprint feature is the dynamic voiceprint feature, obtaining the device state representation, and determining at least one power device state to be detected in the voiceprint segment to be monitored based on the device state representation.

14. The system according to claim 10, wherein Performing time-frequency analysis on the voiceprint segment to be monitored to obtain voiceprint features at multiple resolution levels, including: Performing multi-scale time-frequency analysis on the voiceprint segment to be monitored to obtain the initial voiceprint features at at least one frequency band decomposition level; Based on the frequency band decomposition level, align to the target initial voiceprint features in the initial voiceprint features, and perform time-frequency grid alignment on the target initial voiceprint features to obtain the voiceprint features with alignment completed; Perform time-frequency compounding on the voiceprint features with alignment completed and the installation topology features to obtain compound voiceprint features, where the compound voiceprint features include at least one voiceprint sub-feature; Perform eigen-coupling analysis on the voiceprint sub-features to obtain the voiceprint features with energy ratio completed, and perform dynamic reconstruction on the initial voiceprint features based on the voiceprint features with energy ratio completed to obtain the transient voiceprint features at multiple resolution levels; Based on the frequency band energy levels of the transient voiceprint features, perform an energy level sorting operation on the transient voiceprint features to obtain the first frequency band energy level sequence; Align to the transient voiceprint feature with the lowest frequency band energy level in the transient voiceprint features to obtain the dynamic transient voiceprint feature, and use the dynamic transient voiceprint feature as the first transient coupling mode; Based on the first frequency band energy level sequence, align to the next transient voiceprint feature of the dynamic transient voiceprint feature in the transient voiceprint features to obtain the first coupling mode to be coupled of the dynamic transient voiceprint feature; Perform frequency band dimension reallocation on the dynamic transient voiceprint feature to obtain the reallocated voiceprint features with the preset number of dimensions; Perform energy level gain compensation on the frequency band energy levels of the reallocated voiceprint features to obtain the voiceprint features after energy level gain compensation, where the frequency band energy levels of the voiceprint features after energy level gain compensation are consistent with the frequency band energy levels of the first coupling mode to be coupled; Perform frequency band dimension reallocation on the first coupling mode to be coupled to obtain the reallocated first coupling mode to be coupled with the preset number of dimensions; Perform cross-frequency band compounding on the voiceprint features after energy level gain compensation and the reallocated first coupling mode to be coupled to obtain the second transient coupling mode; Use the first coupling mode to be coupled as the dynamic transient voiceprint feature; Recursively trigger the process of aligning the next transient acoustic feature of the dynamic transient acoustic feature in the transient acoustic features based on the first frequency band energy level sequence, and terminate the recursive process until all transient acoustic features complete cross-level coupling, obtaining at least one transient coupling mode, where the transient coupling mode includes the first transient coupling mode and the second transient coupling mode; Perform an energy level sorting operation on the transient coupling mode based on the frequency band energy level of the transient coupling mode to obtain a second frequency band energy level sequence; Align to the transient coupling mode with the highest frequency band energy level in the transient coupling mode to obtain a dynamic coupling mode, and perform time-frequency compounding on the features of different frequency band dimensions in the dynamic coupling mode to obtain a first coupling mode; Based on the second frequency band energy level sequence, align to the next transient coupling mode of the dynamic coupling mode in the transient coupling mode to obtain a second to-be-coupled mode of the dynamic coupling mode; Perform time-frequency compounding on the dynamic coupling mode and the second to-be-coupled mode to obtain a second coupling mode, and use the second to-be-coupled mode as the dynamic coupling mode; Recursively trigger the process of aligning to the next transient coupling mode of the dynamic coupling mode in the transient coupling mode based on the second frequency band energy level sequence, and terminate the recursive process until all transient coupling modes complete cross-level coupling, obtaining at least one coupling mode, where the coupling mode includes the first coupling mode and the second coupling mode; Use each coupling mode as the acoustic feature of a resolution level.

15. The system according to claim 10, wherein After calculating the coupling coefficient between the acoustic feature and the vibration feature in each time-frequency unit, it further includes: Real-time monitor the current operating environment parameters of the power equipment, where the current operating environment parameters include temperature, humidity, and electromagnetic interference intensity; Based on the preset adjustment rule of the environmental parameter and the coupling coefficient, dynamically adjust the coupling coefficient in each time-frequency unit according to the change of the currently monitored current operating environment parameters.

16. The system according to claim 12, wherein The time-frequency compounding of the reference acoustic-vibration joint feature maps of the same resolution level to obtain the acoustic-vibration joint feature maps of each resolution level includes: Use the reference acoustic-vibration joint feature maps of the same resolution level as input data and input them into a pre-trained deep learning model, where the deep learning model is trained based on the acoustic-vibration joint feature map data with annotations in the normal and abnormal states of the power equipment; Perform feature extraction and fusion processing on the input data through the convolutional layer, pooling layer, and fully connected layer structure of the deep learning model, and output the acoustic-vibration joint feature maps of each resolution level.

17. The system according to claim 13, wherein After performing cross-modal coupling on the transient device state representation, the dynamic acoustic-vibration joint feature map, and the dynamic acoustic feature to obtain the dynamic device state representation, it further includes: Perform reliability verification on the dynamic device state representation, and extract the confirmed accurate device state representation data under the operating conditions similar to the current monitoring period from the historical power equipment operation data as the verification sample set; Use the similarity measurement algorithm to calculate the similarity between the dynamic device state representation and each sample in the verification sample set; If the similarity is lower than a preset reliability threshold, the process of detecting at least one power equipment state to be detected in the to-be-monitored voiceprint segment based on the combined acoustic-vibration feature map and the voiceprint feature is re-executed until the obtained dynamic equipment state characterization passes the reliability verification.

18. The system according to claim 14, wherein The step of performing an energy level sorting operation on the transient coupling mode based on the frequency band energy level of the transient coupling mode to obtain a second frequency band energy level sequence further includes: For each transient coupling mode, according to the corresponding current frequency band energy level, the corresponding feature enhancement coefficient is determined according to a preset mapping relationship between the energy level and the enhancement coefficient; The feature of the transient coupling mode in the frequency band dimension is weighted by using the feature enhancement coefficient.

19. A server system, characterized in that, It includes a server, and the server is configured to execute the method according to any one of claims 1-9.

20. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed, the method according to any one of claims 1-9 is implemented.

Citation Information

Cited By

  • Mine draw shaft state automatic monitoring method based on acoustics and vibration coupling

    CN122388503A