Non-contact power equipment voiceprint monitoring method and system based on deep learning

Through the voiceprint monitoring method of non-contact power equipment based on deep learning, the model and equipment status subdivision type library are enhanced by the pre-trained voiceprint feature, and the safety risks and accuracy problems of traditional monitoring methods are solved, and efficient and accurate monitoring of the operating status of power equipment is achieved.

CN120048286AInactive Publication Date: 2025-05-27BEIJING ZHONGKE DONGREN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192824.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional power equipment monitoring methods have safety risks and are inconvenient to implement on a large scale, and complex on-site environments lead to interference from voiceprint signals, affecting monitoring accuracy.

Method used

The voiceprint monitoring method of non-contact power equipment based on deep learning is adopted. The original voiceprint signal is collected in the preset safety area through the handheld voiceprint acquisition device, and the signal is enhanced by the pre-trained voiceprint feature enhancement model processing, and the device operation status is determined in combination with the device status subdivided type library.

Benefits of technology

It realizes efficient and accurate monitoring of the operating status of power equipment, avoids the safety risks of contact monitoring, and reduces the impact of environmental interference on monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048286A_ABST
    Figure CN120048286A_ABST
Patent Text Reader

Abstract

The invention discloses a non-contact power equipment voiceprint monitoring method and system based on deep learning, and the method comprises the steps: firstly collecting an original voiceprint signal of target power equipment in a safety region through handheld equipment, and then processing the signal through a pre-trained voiceprint feature enhancement model to obtain a current voiceprint, and finally, the final operation state of the target equipment is determined in combination with the equipment state subdivision type library, and efficient and accurate monitoring of the operation state of the power equipment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment maintenance, and more particularly, to a non-contact power equipment voiceprint monitoring method and system based on deep learning. Background Art

[0002] In the operation of the power system, it is crucial to monitor the operation status of power equipment in a timely and accurate manner. Traditional monitoring methods have certain limitations. For example, some methods that require contact with the equipment may pose safety risks and are not convenient for large-scale implementation. With the development of technology, non-contact monitoring has become a trend. As an effective means, voiceprint monitoring can determine the status of power equipment by analyzing the sound characteristics during its operation. However, the complex on-site environment is likely to interfere with the voiceprint signal, affecting the accuracy of monitoring. Summary of the Invention

[0003] The purpose of the present invention is to provide a non-contact power equipment voiceprint monitoring method and system based on deep learning.

[0004] In a first aspect, an embodiment of the present invention provides a non-contact power equipment voiceprint monitoring method based on deep learning, including:

[0005] Collecting the voiceprint of a target power equipment within a preset safe area by using a handheld voiceprint collection device to obtain an original voiceprint signal;

[0006] Inputting the original voiceprint signal into a pre-trained voiceprint feature enhancement model for processing to obtain the current power equipment voiceprint;

[0007] Determining the final operation status of the target power equipment based on the current power equipment voiceprint in combination with the equipment status breakdown type library.

[0008] In a second aspect, an embodiment of the present invention provides a server system, including a server, and the server is used to execute the method described in the first aspect.

[0009] Compared with the prior art, the beneficial effects provided by the present invention include: By using the non-contact power equipment voiceprint monitoring method and system based on deep learning disclosed in the present invention, the original voiceprint signal of the target power equipment is collected in a safe area by using a handheld device, then the signal is processed by using a pre-trained voiceprint feature enhancement model to obtain the current voiceprint, and finally the final operation status of the target equipment is determined in combination with the equipment status breakdown type library, realizing the efficient and accurate monitoring of the operation status of power equipment. Brief Description of the Drawings

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the attached drawings required for the embodiments. It should be understood that the following attached drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related attached drawings can also be obtained based on these attached drawings.

[0011] Figure 1 It is a schematic flowchart of the steps of the non-contact power equipment voiceprint monitoring method based on deep learning provided by the embodiments of the present invention;

[0012] Figure 2 It is a schematic block diagram of the structure of the computer device provided by the embodiments of the present invention. Detailed implementation manners

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the attached drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Usually, the components of the embodiments of the present invention described and shown in the attached drawings here can be arranged and designed in various different configurations.

[0014] The following will, with reference to the attached drawings, detail the specific implementation manners of the present invention.

[0015] To solve the technical problems in the foregoing background technology, Figure 1 It is a schematic flowchart of the non-contact power equipment voiceprint monitoring method based on deep learning provided by the embodiments of the present disclosure. The following will detail the non-contact power equipment voiceprint monitoring method based on deep learning.

[0016] Step S201: Collect the voiceprint of the target power equipment within a preset safe area through a handheld voiceprint collection device to obtain an original voiceprint signal;

[0017] Step S202: Input the original voiceprint signal into a pre-trained voiceprint feature enhancement model for processing to obtain the current power equipment voiceprint;

[0018] Step S203: Based on the current power equipment voiceprint and combined with the equipment status subdivision type library, determine the final equipment operation status corresponding to the target power equipment.

[0019] In an embodiment of the present invention, exemplarily, in a large substation, there are various power equipment operating continuously, such as transformers, switchgear, etc. To ensure the normal operation of these devices and timely detect potential problems, we arrange staff to use handheld voiceprint acquisition devices to perform voiceprint acquisition work. The server sets a preset safety area around the target power equipment. For example, for a large transformer, this safety area may be a range with a radius of 5 meters centered on the transformer (the specific range can be set according to the actual situation and safety specifications). For example, the handheld voiceprint acquisition device supports full-frequency coverage from 20 Hz to 20 kHz, the sampling rate is ≥48 kHz, the dynamic range is ≥96 dB, and it is equipped with a directional microphone (directivity index ≥8 dB) to suppress ambient noise; the preset safety area is dynamically adjusted according to the device type. For example, the default radius of the transformer is 5 meters, and that of the switchgear is 2 meters, and the safety distance rule library is called through the device type recognition module for real-time matching. The staff enters this area with the handheld voiceprint acquisition device and aligns the acquisition end of the device with the operating transformer. At this time, the transformer emits various sounds during operation, and these sounds contain a lot of information about its operating state. The handheld voiceprint acquisition device starts to collect the sound signals emitted by the transformer and converts these sound signals into digital-form original voiceprint signals, and then transmits the collected original voiceprint signals to the server by means of wireless transmission or wired connection. After receiving the original voiceprint signals transmitted from the handheld voiceprint acquisition device, the server starts to process them. For example, the preset standard voiceprint signals stored in the server were collected and processed before when the transformer was in an ideal operating state. The server first obtains the original voiceprint signals and this standard voiceprint signal, and then calculates the time-frequency domain matching degree of the acoustic feature dimensions between them. Suppose the original voiceprint signals deviate from the standard voiceprint signals at some frequencies due to factors such as the on-site environment, and the time-frequency domain matching degree shows that the matching degree in the frequency band of 500 Hz - 1000 Hz is only 60% (just an example to illustrate the matching degree situation). Then, the server performs power frequency phase synchronization processing on the standard voiceprint signal and the original voiceprint signal based on this time-frequency domain matching degree. After processing, the power frequency phase-synchronized standard voiceprint signal and the power frequency phase-synchronized original voiceprint signal are obtained. Then, the server generates a simulation interference feature waveform based on the power frequency phase-synchronized standard voiceprint signal. For example, according to past experience and data analysis, knowing some electromagnetic interference characteristics that may exist at the substation site, a similar interference waveform is simulated. Then, this simulation interference feature waveform is used to perform electromagnetic interference filtering processing on the power frequency phase-synchronized original voiceprint signal to remove those simulated interference components, and the preprocessed voiceprint signal of the original voiceprint signal is obtained. After that, an interference suppression frequency domain matrix for the preprocessed voiceprint signal is generated according to the standard voiceprint signal, the simulation interference feature waveform, the original voiceprint signal, and the preprocessed voiceprint signal.Based on this interference suppression frequency-domain matrix, perform frequency-domain feature reconstruction on the preprocessed voiceprint signal to obtain a voiceprint signal with electromagnetic interference filtered out. The server then decomposes the voiceprint signal with electromagnetic interference filtered out from the transient feature dimension power frequency harmonics to the power frequency band dimension, obtaining the harmonic component data of the power frequency band dimension of the voiceprint signal with electromagnetic interference filtered out, which includes amplitude spectrum components and phase spectrum components. Suppose the value of the amplitude spectrum component at a certain frequency point after decomposition is 0.8 (just an example to illustrate the specific value), and the phase spectrum component is 30 degrees (also an example). Then, perform adaptive parameter estimation and optimization processing on the amplitude spectrum component, and perform the same processing on the phase spectrum component. For example, continuously adjust the parameters through some intelligent algorithms to make these components more accurately reflect the characteristics of the voiceprint signal. After processing, obtain the amplitude spectrum component and phase spectrum component after adaptive parameter estimation and optimization, and then determine the suppression feature signal in the power frequency band dimension based on them. At the same time, the server extracts the acoustic feature dimension data of the voiceprint signal with electromagnetic interference filtered out in the transient feature dimension based on the voiceprint feature enhancement model, and then generates a transient interference suppression matrix for these acoustic feature dimension data based on the model. Use this matrix to perform time-domain pulse cancellation on the acoustic feature dimension data to obtain the suppression feature signal in the transient feature dimension. After that, generate the first voiceprint feature reconstruction coefficient of the suppression feature signal in the power frequency band dimension and the second voiceprint feature reconstruction coefficient of the suppression feature signal in the transient feature dimension based on the voiceprint feature enhancement model, and use these two coefficients to perform feature space synthesis on the suppression feature signal in the power frequency band dimension and the suppression feature signal in the transient feature dimension to obtain a voiceprint signal with reduced power frequency noise. Finally, obtain the feature stability compensation parameter for the voiceprint feature energy based on the voiceprint feature enhancement model, and use this parameter to perform voiceprint feature reconstruction on the voiceprint signal with reduced power frequency noise, so as to obtain the voiceprint of the current power equipment (here it is the transformer collected before), that is, the voiceprint signal that can more accurately reflect the equipment operation state after a series of optimization processes. After the server obtains the voiceprint of the current power equipment, it starts to combine the equipment status subdivision type library to determine the final equipment operation state of the target power equipment (still taking the transformer as an example). First, the server performs voiceprint feature characterization on the voiceprint of the current power equipment to obtain the target voiceprint features. Then, based on this target voiceprint feature, obtain multiple pending equipment statuses from the equipment status subdivision type library. This equipment status subdivision type library is a data set of multiple subdivision operation statuses of the same equipment operation state (such as the operation state of a transformer). For example, for the operation state of a transformer, the subdivision operation states may include normal operation but slightly overloaded, normal operation and load balanced, and there are slight potential fault hazards during operation (such as a slight looseness of a certain component). The server obtains the typical working condition voiceprint sample sets corresponding to these multiple equipment status subdivision types in the equipment status subdivision type library, and each equipment status subdivision type corresponds to at least one typical voiceprint sample in the typical working condition voiceprint sample set.Suppose there are 5 typical voiceprint samples for the sub - state of "normal operation but slightly overloaded". The server performs voiceprint feature characterization on each typical voiceprint sample in this set of typical voiceprint samples to obtain corresponding typical voiceprint features. Then, it classifies the obtained multiple typical voiceprint features to obtain voiceprint feature clusters corresponding to each sub - type of device state. A voiceprint feature cluster contains at least one typical voiceprint feature, and at least one typical voiceprint feature contains the centroid of the voiceprint features of this voiceprint feature cluster. The server takes the centroid of the voiceprint features of the voiceprint feature cluster corresponding to each sub - type of device state as the reference voiceprint feature of each sub - type of device state, and then determines the first voiceprint matching degree between the target voiceprint feature and the reference voiceprint feature of each sub - type of device state. For example, the first voiceprint matching degree between the target voiceprint feature and the reference voiceprint feature of the sub - state of "normal operation but slightly overloaded" is 70% (just an example to illustrate the matching degree value). Then, it sorts the multiple first voiceprint matching degrees in descending order of matching degree priority, and filters out the sub - types of device states corresponding to the preset number of first voiceprint matching degrees with higher matching degrees as the pending device states. Suppose the preset number is 3, then the 3 sub - states with the highest matching degrees are filtered out as the pending device states. After that, based on the voiceprint feature descriptions and key acoustic dimensions corresponding to these pending device states respectively, a first voiceprint feature parsing instruction is generated. This key acoustic dimension is obtained through a certain method. For example, taking the second voiceprint feature parsing instruction as a control parameter, loading it into the acoustic feature parsing unit for acoustic feature space mapping to obtain an acoustic reference parameter set for identifying different sub - types of device states, and obtaining the key acoustic dimension from this acoustic reference parameter set. Through the acoustic feature parsing unit, based on the voiceprint feature correlation parameters included in the first voiceprint feature parsing instruction, the first voiceprint feature parsing instruction is parsed for voiceprint features to obtain a first feature parsing output, which contains the first voiceprint feature matching strategy for the key acoustic dimensions corresponding to multiple pending device states respectively. Then, the acoustic feature parsing unit executes the first voiceprint feature matching strategy to obtain the feature parameters of multiple pending device states respectively under the key acoustic dimension from the voiceprint feature database. Finally, based on the current power equipment voiceprint, the voiceprint feature descriptions corresponding to multiple pending device states respectively, and the feature parameters of multiple pending device states respectively under the key acoustic dimension, the final device operation state of the current power equipment voiceprint is obtained from multiple pending device states. For example, by generating a composite voiceprint analysis instruction according to the preset voiceprint feature fusion rule, and performing acoustic feature modeling operations on the acoustic parameter descriptions in the instruction through the voiceprint feature analysis model, it is finally determined that the transformer is in the final device operation state of "normal operation and load balancing".

[0020] In an embodiment of the present invention, the final device operating state corresponding to the target power device determined based on the current power device voiceprint and the device state subdivision type library can be implemented through the following examples.

[0021] Perform voiceprint feature characterization on the current power device voiceprint to obtain target voiceprint features;

[0022] Based on the target voiceprint features, obtain multiple pending device states from the device state subdivision type library, where the device state subdivision type library is a data set of multiple subdivision operating states of the same device operating state;

[0023] Based on the voiceprint feature descriptions respectively corresponding to the multiple pending device states, identify the characteristic parameters of the multiple pending device states in the key acoustic dimensions; wherein, the key acoustic dimension is the voiceprint characteristic dimension used to identify the multiple pending device states;

[0024] Based on the current power device voiceprint, the voiceprint feature descriptions respectively corresponding to the multiple pending device states, and the characteristic parameters of the multiple pending device states in the key acoustic dimensions, obtain the final device operating state of the current power device voiceprint from the multiple pending device states.

[0025] In an embodiment of the present invention, for example, after the server receives the voiceprint of the current power device (such as a certain transformer in a substation), it starts the voiceprint feature characterization work. The server takes the just obtained target voiceprint features and goes to the device state subdivision type library to find matching information. The server compares the target voiceprint features with the voiceprint features corresponding to these subdivision states, finds those that are relatively similar, and determines these corresponding subdivision states as multiple pending device states. For each pending device state, the server deeply explores according to its corresponding voiceprint feature description. For example, for the pending state of "normal operation and slightly overloaded", its voiceprint feature description may involve the change characteristics of the sound intensity in certain frequency bands, etc. The server uses specific acoustic analysis techniques to accurately identify the specific characteristic parameters of each pending device state in these key dimensions (such as a specific frequency range, the harmonic components of the sound, etc.) along the key acoustic dimensions used to identify these pending states, such as the amplitude size at a certain key frequency. The server integrates the current power device voiceprint, the voiceprint feature descriptions corresponding to each pending device state, and their characteristic parameters in the key acoustic dimensions. Then, through a carefully designed determination mechanism, which may involve complex data analysis, comparison, and weight calculation, etc., finally determine the most accurate final device operating state corresponding to the current power device voiceprint from these pending device states. For example, it is determined that the transformer is in a state of normal operation and slightly high temperature at this time.

[0026] In an embodiment of the present invention, based on the voiceprint feature descriptions respectively corresponding to the multiple pending device states, identifying the characteristic parameters of the multiple pending device states respectively in the key acoustic dimensions can be implemented through the following examples.

[0027] Generate a first voiceprint feature parsing instruction based on the voiceprint feature descriptions respectively corresponding to the multiple pending device states and the key acoustic dimensions;

[0028] Using the first voiceprint feature parsing instruction as a control parameter, load it into the acoustic feature parsing unit in a pre-trained voiceprint feature analysis model for acoustic feature space mapping, and obtain the characteristic parameters of the multiple pending device states respectively in the key acoustic dimensions.

[0029] In an embodiment of the present invention, for example, after the server determines multiple pending device states (such as "normal operation and slight overload", "normal operation and slightly high temperature", etc. for a certain transformer in a substation), it will carry out work according to the voiceprint feature description corresponding to each pending device state. These voiceprint feature descriptions contain information such as the intensity change of specific frequency sounds and harmonic characteristics. At the same time, in combination with the key acoustic dimensions used to identify these pending states (such as specific frequency ranges, phase characteristics of sounds, etc.), the server integrates and encodes this information according to a pre-set instruction generation rule to generate a first voiceprint feature parsing instruction. The server takes the generated first voiceprint feature parsing instruction and uses it as a key control parameter to send it to the acoustic feature parsing unit in a pre-trained voiceprint feature analysis model. After operation, finally, the specific characteristic parameters of each pending device state respectively in the key acoustic dimensions are accurately obtained, such as the exact amplitude, phase, etc. values at specific key frequencies.

[0030] In an embodiment of the present invention, the step of using the first voiceprint feature parsing instruction as a control parameter, loading it into the acoustic feature parsing unit in a pre-trained voiceprint feature analysis model for acoustic feature space mapping, and obtaining the characteristic parameters of the multiple pending device states respectively in the key acoustic dimensions can be implemented through the following examples.

[0031] Through the acoustic feature parsing unit, based on the voiceprint feature correlation parameters included in the first voiceprint feature parsing instruction, perform voiceprint feature parsing on the first voiceprint feature parsing instruction to obtain a first feature parsing output; the first feature parsing output includes a first voiceprint feature matching strategy for the key acoustic dimensions respectively corresponding to the multiple pending device states;

[0032] Execute the first voiceprint feature matching strategy through the acoustic feature parsing unit to obtain the characteristic parameters of the multiple pending device states respectively in the key acoustic dimensions from the voiceprint feature database.

[0033] In an embodiment of the present invention, exemplarily, after the server transmits the first voiceprint feature parsing instruction to the acoustic feature parsing unit in the pre-trained voiceprint feature analysis model, the unit starts to work. It first carefully analyzes the voiceprint feature correlation parameters included in the first voiceprint feature parsing instruction. These parameters are like keys that can open the door to in-depth understanding of the instruction. For example, the correlation parameters may involve information such as the voiceprint feature weights of the to-be-determined device states in a specific frequency range and the correlation relationships between different acoustic dimensions. Based on these parameters, the acoustic feature parsing unit uses its built-in complex algorithms and models to comprehensively and meticulously parse the voiceprint features of the first voiceprint feature parsing instruction. After a series of precise calculations, a first feature parsing output is finally obtained. This output includes the first voiceprint feature matching strategies corresponding to the key acoustic dimensions for each to-be-determined device state (such as different operation sub-states of a substation transformer). These strategies clearly stipulate how to compare and match voiceprint features in different key acoustic dimensions. After obtaining the first voiceprint feature matching strategies in the first feature parsing output, the acoustic feature parsing unit immediately takes action. According to these matching strategies, it searches for the required information in the voiceprint feature database. The acoustic feature parsing unit accurately screens and extracts the feature parameters of each to-be-determined device state in the key acoustic dimensions from the database according to the requirements of the matching strategies for each to-be-determined device state in a specific key acoustic dimension, such as the amplitude matching rule at a certain key frequency. For example, it can accurately obtain the specific amplitude, phase and other feature parameter values of a to-be-determined state at a specific key frequency.

[0034] In an embodiment of the present invention, the key acoustic dimensions are obtained in the following manner.

[0035] The key acoustic dimensions are obtained by retrieving the industry standard voiceprint feature library;

[0036] Or,

[0037] Taking the second voiceprint feature parsing instruction as a control parameter, it is loaded into the acoustic feature parsing unit for acoustic feature space mapping to obtain an acoustic reference parameter set for identifying different device state sub-types, where the second voiceprint feature parsing instruction includes the following relevant feature dimensions in the voiceprint feature description: the voiceprint feature description of the same device operation state to which the device state sub-type library belongs, the voiceprint feature descriptions of multiple device state sub-types in the device state sub-type library;

[0038] The key acoustic dimensions are obtained from the acoustic reference parameter set.

[0039] In an embodiment of the present invention, exemplarily, when the server needs to determine the key acoustic dimensions, it first interacts with an industry standard acoustic feature library that stores a large amount of industry standard information. This library details the widely recognized acoustic feature-related standards for various devices in different operating states. The server accurately searches in the library for the acoustic feature standard content related to the currently monitored power device (such as a transformer in a substation) according to the established query rules and procedures. It directly extracts from this the key acoustic dimension information used to accurately identify different sub-types of device states. For example, for a transformer, it may obtain features such as the frequency change pattern of the sound within a specific frequency band and the characteristics of certain harmonic components as key acoustic dimensions. These dimensions are acoustic feature dimensions that have been determined through long-term industry practice and research and are crucial for judging the device state. The server generates a second acoustic feature parsing instruction according to the requirements. This instruction integrates the acoustic feature descriptions of the same device operating state to which the sub-type library of device states belongs, as well as the relevant feature dimension information in the acoustic feature descriptions of multiple sub-types of device states in the library. Then, the server transmits the second acoustic feature parsing instruction as an important control parameter to the acoustic feature parsing unit. After receiving the instruction, this unit embarks on a "journey of exploration" in the acoustic feature space. It performs detailed spatial mapping operations on the relevant acoustic data according to the requirements in the instruction, using complex algorithms and models, to obtain an acoustic reference parameter set for identifying different sub-types of device states. This acoustic reference parameter set contains a large number of acoustic feature data related to different sub-types of device states. The server then precisely identifies from this set the key acoustic dimensions that can effectively distinguish different sub-types of device states through a specific screening and identification mechanism. For example, it determines certain specific frequency ranges, specific phase characteristics of the sound, etc. as key acoustic dimensions, so as to more accurately judge the actual operating state of the device based on these dimensions in the subsequent process.

[0040] In an embodiment of the present invention, taking the second acoustic feature parsing instruction as a control parameter and loading it into the acoustic feature parsing unit for acoustic feature space mapping to obtain an acoustic reference parameter set for identifying different sub-types of device states can be implemented through the following example.

[0041] Through the acoustic feature parsing unit, based on the acoustic feature correlation parameters included in the second acoustic feature parsing instruction, perform acoustic feature parsing on the second acoustic feature parsing instruction to obtain a second feature parsing output; the second feature parsing output includes a second acoustic feature matching strategy for identifying different sub-types of device states;

[0042] Execute the second acoustic feature matching strategy through the acoustic feature parsing unit to obtain an acoustic reference parameter set for identifying different sub-types of device states.

[0043] In an embodiment of the present invention, exemplarily, after receiving the second voiceprint feature parsing instruction, the server transmits it to the acoustic feature parsing unit. This unit immediately starts processing. It first focuses on the voiceprint feature correlation parameters included in the second voiceprint feature parsing instruction. For example, these correlation parameters may involve important information such as the weights of voiceprint features of various types in the device state subdivision type library in a specific frequency band, the correlation relationship between different acoustic dimensions, etc. Based on these parameters, the acoustic feature parsing unit uses its built-in professional algorithms and analysis models to conduct a comprehensive and detailed voiceprint feature parsing of the second voiceprint feature parsing instruction. After a series of complex operations and processing, a second feature parsing output is finally obtained. This output includes the second voiceprint feature matching strategies for identifying different device state subdivision types. These strategies clearly stipulate how to perform matching and judgment based on specific acoustic features when facing different device state subdivision types. After receiving the second voiceprint feature matching strategies in the second feature parsing output, the acoustic feature parsing unit immediately starts to execute. It will, based on these matching strategies, on the basis of a large amount of acoustic data and related analysis models it has mastered, conduct one-by-one comparison and analysis of different device state subdivision types. Through such rigorous and detailed operations, an acoustic reference parameter set for identifying different device state subdivision types is finally successfully obtained. This set contains a lot of acoustic feature data closely related to different device state subdivision types, providing an important basis for accurately judging the device state subsequently.

[0044] In an embodiment of the present invention, obtaining multiple pending device states from the device state subdivision type library based on the target voiceprint feature can be implemented through the following examples.

[0045] Obtain the reference voiceprint features of each device state subdivision type in the device state subdivision type library, and determine the first voiceprint matching degrees between the target voiceprint feature and the reference voiceprint features of each device state subdivision type;

[0046] Sort the multiple first voiceprint matching degrees in descending order of matching degree priority, and screen out the device state subdivision types corresponding to a preset number of the first voiceprint matching degrees with higher voiceprint matching degrees as the pending device states.

[0047] In an embodiment of the present invention, exemplarily, after receiving a target voiceprint feature (such as a feature obtained after processing the voiceprint of a certain power equipment in a substation), the server starts to interact with the equipment status subdivision type library. The server obtains the reference voiceprint features of these various equipment status subdivision types one by one. Then, the server uses a professional voiceprint matching algorithm to carefully compare the target voiceprint feature with the reference voiceprint feature of each equipment status subdivision type. For example, for the subdivision states of a transformer such as "normal operation and slightly overloaded" and "normal operation and slightly higher temperature", the similarity degrees of the target voiceprint feature and their respective reference voiceprint features in terms of frequency, amplitude, phase, etc. are calculated respectively, so as to determine the first voiceprint matching degrees between the target voiceprint feature and the reference voiceprint features of each equipment status subdivision type. After obtaining all the first voiceprint matching degrees, the server sorts them in descending order of matching degree. Then, according to a preset quantity (such as preset to 3), the equipment status subdivision types corresponding to the top-ranked voiceprint matching degrees are selected from the sorted results. For example, the equipment status subdivision types such as "normal operation and slightly overloaded" and "normal operation with a slight abnormal sound in a certain component" corresponding to the top 3 matching degrees will be selected as the pending equipment statuses for further analysis and judgment of which specific operating state the power equipment is in.

[0048] In an embodiment of the present invention, the obtaining of the reference voiceprint features of the various equipment status subdivision types in the equipment status subdivision type library can be implemented through the following examples.

[0049] Obtain the typical working condition voiceprint sample sets corresponding to multiple equipment status subdivision types in the equipment status subdivision type library, and each equipment status subdivision type corresponds to at least one typical voiceprint sample in the typical working condition voiceprint sample set;

[0050] Perform voiceprint feature characterization on each typical voiceprint sample in the typical working condition voiceprint sample set to obtain the corresponding typical voiceprint features;

[0051] Perform voiceprint pattern classification on the obtained multiple typical voiceprint features to obtain the voiceprint feature clusters corresponding to the respective multiple equipment status subdivision types, where a voiceprint feature cluster contains at least one typical voiceprint feature, and the at least one typical voiceprint feature contains the voiceprint feature centroid of the voiceprint feature cluster;

[0052] Take the voiceprint feature centroids of the voiceprint feature clusters corresponding to each equipment status subdivision type as the reference voiceprint features of the respective equipment status subdivision types.

[0053] In an embodiment of the present invention, exemplarily, the server first establishes a connection with the device status subdivision type library. For the target power device to be monitored (such as a transformer in a substation), the library stores relevant data for different device status subdivision types. The server obtains the typical working condition voiceprint sample sets corresponding to each device status subdivision type from it. For example, for the subdivision type of "normal operation and slight overload" of the transformer, the library stores multiple voiceprint sample sets under multiple typical working conditions collected when this state occurred many times in the past. Each sample set contains several voiceprint samples collected under the corresponding working conditions. After the server obtains these typical working condition voiceprint sample sets, it begins to perform fine processing on each typical voiceprint sample among them. It uses a professional voiceprint feature characterization algorithm. For each typical voiceprint sample, the server will analyze its feature performance in multiple acoustic dimensions such as frequency, amplitude, and phase, and extract the typical voiceprint features that can accurately describe this sample. For example, features such as the energy distribution in a specific frequency band and the slope of frequency change of a certain typical voiceprint sample are accurately extracted. The server sorts out the numerous typical voiceprint features extracted from each typical working condition voiceprint sample set. Then, according to a specific voiceprint pattern classification algorithm, these typical voiceprint features are classified according to the device status subdivision types to which they belong. Taking the transformer as an example, for different subdivision types such as "normal operation and slight overload" and "normal operation and slightly high temperature", the server will classify the typical voiceprint features belonging to the same subdivision type into one category, forming their corresponding voiceprint feature clusters. Each cluster contains at least one typical voiceprint feature, and there is a typical voiceprint feature that can represent the core feature of this cluster, that is, the voiceprint feature centroid. After the server completes the division of the voiceprint feature clusters, it extracts the voiceprint feature centroids in the voiceprint feature clusters corresponding to each device status subdivision type. These voiceprint feature centroids are the "standard templates" of each device status subdivision type and can accurately represent the typical voiceprint features of this subdivision type. The server takes them as the reference voiceprint features of each device status subdivision type respectively, so as to be used as an important reference basis for subsequent operations such as comparing with the target voiceprint features.

[0054] In an embodiment of the present invention, based on the target voiceprint feature, obtaining multiple pending device statuses from the device status subdivision type library can be implemented through the following examples.

[0055] Perform acoustic feature modeling on the voiceprint feature descriptions of each device status subdivision type in the device status subdivision type library to obtain corresponding acoustic feature templates;

[0056] Determine the second voiceprint matching degrees between the target voiceprint feature and the acoustic feature templates of each device status subdivision type;

[0057] Sort the multiple second voiceprint matching degrees in descending order of matching degree priority, and screen out the preset number of device status subdivision types corresponding to the second voiceprint matching degrees with the highest ranking as the pending device status.

[0058] In an embodiment of the present invention, for example, the server accesses the device status subdivision type library, which stores detailed voiceprint feature description information of various device status subdivision types (such as the normal operation, minor faults, etc. of a certain power device in a substation). For the voiceprint feature description of each device status subdivision type, the server uses professional acoustic analysis algorithms and modeling tools to perform acoustic feature modeling. By analyzing key elements such as frequency range, amplitude change law, and harmonic components in the voiceprint feature description, an exclusive acoustic feature template is created for each device status subdivision type, which accurately depicts the voiceprint feature pattern in each subdivision state. After the server obtains the generated target voiceprint feature (from the result of processing the voiceprint of the actual monitored power device), it compares it one by one with the acoustic feature templates of each device status subdivision type created before. Using professional matching algorithms, the similarity between the two is considered from multiple aspects such as frequency, amplitude, and phase, so as to determine the second voiceprint matching degree between the target voiceprint feature and the acoustic feature template of each device status subdivision type. For example, for the acoustic feature template of the "normal operation and slightly overloaded" subdivision state of a certain power device, the specific matching degree value with the target voiceprint feature is calculated through detailed comparison. After the server obtains all the second voiceprint matching degrees, it sorts them in descending order of matching degree priority. Just like ranking these matching degrees, the higher the matching degree, the more similar the target voiceprint feature is to the acoustic feature template of this device status subdivision type. Then, according to the preset number (such as preset to 3), the device status subdivision types corresponding to the higher-ranking matching degrees are screened out from the sorted results. For example, the device status subdivision types such as "normal operation and slightly overloaded" and "normal operation and slight abnormal sound of a certain component" corresponding to the top 3 matching degrees will be screened out as the pending device status for further analysis and judgment of the specific operation status of the power device.

[0059] In an embodiment of the present invention, based on the current power device voiceprint, the voiceprint feature descriptions respectively corresponding to the multiple pending device statuses, and the characteristic parameters of the multiple pending device statuses in the key acoustic dimensions, to obtain the final device operation status of the current power device voiceprint from the multiple pending device statuses, the following example can be executed.

[0060] Generate a composite voiceprint analysis instruction according to the preset voiceprint feature fusion rule, based on the current power device voiceprint, the voiceprint feature descriptions respectively corresponding to the multiple pending device statuses, and the characteristic parameters of the multiple pending device statuses in the key acoustic dimensions;

[0061] Using the composite voiceprint analysis instruction as a control parameter, load it into the voiceprint feature analysis model for acoustic feature space mapping to obtain the final device operating state of the current power device voiceprint.

[0062] In an embodiment of the present invention, exemplarily, after the server obtains the current power device voiceprint (such as the processed voiceprint from a certain transformer in a substation), and the voiceprint feature descriptions corresponding to multiple pending device states (such as "normal operation and slightly overloaded", "normal operation and slightly high temperature", etc.) and their feature parameters in key acoustic dimensions, it will work according to the preset voiceprint feature fusion rules. The server will integrate and arrange the acoustic features of the current power device voiceprint, the voiceprint feature description content of each pending device state (such as the sound performance in a specific frequency band), and their feature parameters in key acoustic dimensions (such as the amplitude at a key frequency) according to specific order, weight and other rules, so as to generate a detailed composite voiceprint analysis instruction. This instruction clearly stipulates the various conditions and requirements for subsequent analysis and processing. The server takes the generated composite voiceprint analysis instruction as a key control parameter and transmits it to the voiceprint feature analysis model. After receiving the instruction, this model starts to work. It will perform complex mapping operations in its internal acoustic feature space according to the requirements in the composite voiceprint analysis instruction. Just like exploring the final answer along the path indicated by the instruction in a huge acoustic maze. By deeply analyzing and integrating the acoustic feature information of the current power device voiceprint and related pending device states, the final device operating state of the current power device voiceprint is finally accurately obtained. For example, it is determined that the transformer is in the specific operating state of "normal operation and slightly overloaded", providing a precise basis for subsequent equipment maintenance, monitoring and other work.

[0063] In an embodiment of the present invention, the step of using the composite voiceprint analysis instruction as a control parameter, loading it into the voiceprint feature analysis model for acoustic feature space mapping to obtain the final device operating state of the current power device voiceprint can be implemented through the following examples.

[0064] Execute the following steps through the voiceprint feature analysis model:

[0065] Perform acoustic feature modeling on the acoustic parameter descriptions in the composite voiceprint analysis instruction to obtain acoustic analysis features, where the acoustic parameter descriptions include: the voiceprint feature descriptions corresponding to the multiple pending device states and the feature parameters of the multiple pending device states in key acoustic dimensions respectively;

[0066] Characterize the current power equipment voiceprint in the composite voiceprint analysis instruction to obtain the target voiceprint feature;

[0067] Perform multi-dimensional integration on the acoustic analysis feature and the target voiceprint feature to obtain the composite acoustic feature;

[0068] Perform multi-stage state decision optimization according to the composite acoustic feature to obtain the final device operation state of the current power equipment voiceprint. Among them, when making a state determination, according to the composite acoustic feature and the confidence evaluation result of the previous stage, obtain the confidence evaluation result of the current stage state determination process.

[0069] In an embodiment of the present invention, exemplarily, after the server transmits the composite voiceprint analysis instruction to the voiceprint feature analysis model, the model first focuses on the acoustic parameter description part in the instruction. The acoustic parameter description here covers the voiceprint feature descriptions corresponding to multiple pending device states (such as the detailed states of "normal operation and slightly overloaded" and "normal operation and slightly high temperature" of a certain power device in a substation) and their respective characteristic parameters in key acoustic dimensions. It will deeply analyze the characteristic manifestations of each pending device state in specific key acoustic dimensions, such as the amplitude and phase changes at a certain key frequency. Through a series of delicate operations and processes, these information will be transformed into acoustic analysis features that can be more intuitively analyzed. At the same time, the model conducts voiceprint feature characterization work on the current power device voiceprint in the composite voiceprint analysis instruction. By analyzing the performance of the voiceprint in multiple aspects such as frequency, amplitude, and phase, specific algorithms are used to extract the features that can uniquely identify the voiceprint, thereby obtaining the target voiceprint features. After obtaining the acoustic analysis features and the target voiceprint features respectively, what the model needs to do next is to integrate them multi-dimensionally. This process is like skillfully splicing and integrating the previously depicted local picture and the carved artwork. It will match, superimpose, and adjust the acoustic analysis features and the target voiceprint features in multiple dimensions according to the preset integration rules and algorithms, so that they can be organically combined together to form a composite acoustic feature containing more comprehensive and accurate acoustic information. Finally, the model conducts multi-stage state decision optimization based on the obtained composite acoustic feature. The model will refer to the confidence evaluation results of the previous stage, such as the credibility obtained when initially analyzing certain features before. Then, combined with the specific situation of the current composite acoustic feature, through a series of complex decision algorithms and optimization mechanisms, a detailed comparison and trade-off are made on various device operation states that the current power device voiceprint may correspond to. After multiple rounds of analysis and judgment, the final device operation state of the current power device voiceprint is accurately obtained. For example, it is determined that the power device in the substation is in the state of "normal operation and slightly overloaded" at this time, and at the same time, the confidence evaluation result of the current stage state determination process can also be obtained, such as giving a credibility value of 80%, indicating the reliability of this determination, providing a precise and valuable reference basis for subsequent device maintenance and monitoring and other work.

[0070] In an embodiment of the present invention, the step of inputting the original voiceprint signal into a pre-trained voiceprint feature enhancement model for processing to obtain the current power device voiceprint can be implemented through the following examples.

[0071] Obtain the original voiceprint signal;

[0072] Based on the voiceprint feature enhancement model, perform electromagnetic interference filtering processing on the original voiceprint signal to obtain the voiceprint signal with electromagnetic interference filtered of the original voiceprint signal;

[0073] Perform power frequency interference elimination processing on the voiceprint signal with electromagnetic interference filtered based on the voiceprint feature enhancement model to obtain a voiceprint signal with reduced power frequency noise of the original voiceprint signal;

[0074] Perform voiceprint feature reconstruction processing on the voiceprint signal with reduced power frequency noise based on the voiceprint feature enhancement model to obtain the current power equipment voiceprint of the original voiceprint signal.

[0075] In an embodiment of the present invention, by way of example, the server remains connected to the handheld voiceprint acquisition device. When the staff completes voiceprint acquisition of the target power equipment (such as the transformer in a substation) within the preset safe area, the acquisition device will transmit the obtained original voiceprint signal to the server. These original voiceprint signals contain various original information of the sound emitted during the operation of the equipment, but may also be mixed with many interference factors. After receiving the original voiceprint signal, the server immediately sends it into the pre-trained voiceprint feature enhancement model. The model first performs electromagnetic interference filtering processing on the original voiceprint signal. In a power environment such as a substation, there are many electromagnetic interference sources, and these interferences will affect the accuracy of the voiceprint signal. The model identifies and analyzes the electromagnetic interference characteristics in the original voiceprint signal, and uses specific algorithms and parameter settings to remove these interference components from the original voiceprint signal, thereby obtaining a voiceprint signal with electromagnetic interference filtered, which more purely reflects the sound characteristics emitted by the equipment itself. Then, the voiceprint feature enhancement model continues to perform power frequency interference elimination processing on the voiceprint signal with electromagnetic interference filtered. When the power equipment operates, power frequency signals will be generated, and their related noises will also be mixed into the voiceprint signal. The model accurately analyzes the power frequency-related interference characteristics in the voiceprint signal with electromagnetic interference filtered, and uses targeted algorithms to adjust parameters such as the frequency and amplitude of the signal, gradually eliminating the influence of power frequency noise, and further obtaining a voiceprint signal with reduced power frequency noise, improving the quality of the voiceprint signal. Finally, the voiceprint feature enhancement model performs voiceprint feature reconstruction processing on the voiceprint signal with reduced power frequency noise. It will reshape and adjust the voiceprint signal in terms of frequency, amplitude, phase, etc. according to the acoustic feature information retained and optimized during the previous processing, so that the finally obtained current power equipment voiceprint can more accurately and comprehensively reflect the actual operating state of the target power equipment.

[0076] In an embodiment of the present invention, the electromagnetic interference filtering processing of the original voiceprint signal based on the voiceprint feature enhancement model to obtain the voiceprint signal with electromagnetic interference filtered of the original voiceprint signal can be implemented through the following examples.

[0077] Obtain a preset standard voiceprint signal;

[0078] Based on the voiceprint feature enhancement model, the original voiceprint signal is subjected to electromagnetic interference filtering processing based on the standard voiceprint signal to obtain the voiceprint signal after electromagnetic interference filtering.

[0079] In an embodiment of the present invention, exemplarily, preset standard voiceprint signals are stored in the server. These standard voiceprint signals are collected and finely processed under specific environments without obvious interference factors when the target power equipment (such as a certain transformer in a substation) is in an ideal operating state. It accurately reflects the voiceprint characteristics that the equipment should emit during normal operation, providing a reliable reference benchmark for subsequent processing of the original voiceprint signal. When the server receives the original voiceprint signal transmitted from the handheld voiceprint acquisition device, it immediately calls the voiceprint feature enhancement model to carry out electromagnetic interference filtering processing. The model uses the obtained standard voiceprint signal as an important reference. Since the standard voiceprint signal represents the ideal voiceprint state of the equipment without interference, the model analyzes which parts of the original voiceprint signal are significantly different from the standard voiceprint signal by comparing the original voiceprint signal with the standard voiceprint signal, and these differences are likely caused by electromagnetic interference. Then, the model uses its built-in professional algorithms and parameter settings to specifically adjust and eliminate the parts of the original voiceprint signal that are suspected of being affected by electromagnetic interference. For example, if it is found that the energy distribution of the original voiceprint signal in a certain frequency band deviates greatly from the standard voiceprint signal, the model will use a specific algorithm to reduce the interference component in this frequency band to make it closer to the characteristics of the standard voiceprint signal in this frequency band. After such a series of meticulous processing, the electromagnetic interference components in the original voiceprint signal are effectively filtered, thus obtaining the voiceprint signal after electromagnetic interference filtering, which can more purely reflect the voiceprint characteristics actually emitted by the target power equipment, laying a good foundation for subsequent further processing and accurate judgment of the equipment operating state.

[0080] In an embodiment of the present invention, the step of, based on the voiceprint feature enhancement model, performing electromagnetic interference filtering processing on the original voiceprint signal based on the standard voiceprint signal to obtain the voiceprint signal after electromagnetic interference filtering can be implemented through the following examples.

[0081] Obtain the time-frequency domain matching degree of the acoustic feature dimension between the standard voiceprint signal and the original voiceprint signal;

[0082] Based on the time-frequency domain matching degree, perform power frequency phase synchronization processing on the standard voiceprint signal and the original voiceprint signal to obtain the power frequency phase synchronized standard voiceprint signal and the power frequency phase synchronized original voiceprint signal;

[0083] Based on the power frequency phase synchronized standard voiceprint signal and the power frequency phase synchronized original voiceprint signal, perform electromagnetic interference filtering processing on the power frequency phase synchronized original voiceprint signal to obtain the voiceprint signal after electromagnetic interference filtering.

[0084] In an embodiment of the present invention, exemplarily, after the server receives the original voiceprint signal and the stored standard voiceprint signal (such as the standard voiceprint signal collected under the ideal operating state of a certain transformer in a substation), it will obtain the time-frequency domain matching degree between them in the acoustic feature dimension through the professional analysis module in the voiceprint feature enhancement model. For example, in the frequency band of 500 Hz - 1000 Hz and a specific time period, the time-frequency domain matching degree between the original voiceprint signal and the standard voiceprint signal may be 70%. Based on the obtained time-frequency domain matching degree, the voiceprint feature enhancement model starts to perform power frequency phase synchronization processing on the standard voiceprint signal and the original voiceprint signal. The standard voiceprint signal and the original voiceprint signal are like two musical melodies with inconsistent rhythms. The model makes their rhythms synchronous by adjusting their phases in the power frequency (related to the operating frequency of electrical equipment). For example, by finely tuning the phase of the original voiceprint signal at certain moments to make it consistent with the phase characteristics of the standard voiceprint signal related to the power frequency, a power frequency phase-synchronized standard voiceprint signal and a power frequency phase-synchronized original voiceprint signal are obtained. After obtaining the power frequency phase-synchronized standard voiceprint signal and the original voiceprint signal, the model further performs electromagnetic interference filtering processing on the power frequency phase-synchronized original voiceprint signal. At this time, since phase synchronization has been achieved, the model can more accurately compare the differences between the two, and filter out the parts of the power frequency phase-synchronized original voiceprint signal that are significantly different from the power frequency phase-synchronized standard voiceprint signal (very likely caused by electromagnetic interference). Thus, a voiceprint signal with electromagnetic interference filtered out is obtained, which can more purely reflect the voiceprint characteristics actually emitted by the device.

[0085] In an embodiment of the present invention, the electromagnetic interference filtering process for the power frequency phase-synchronized original voiceprint signal based on the power frequency phase-synchronized standard voiceprint signal and the power frequency phase-synchronized original voiceprint signal to obtain the voiceprint signal with electromagnetic interference filtered out can be implemented through the following examples.

[0086] Generate a simulated interference feature waveform based on the power frequency phase-synchronized standard voiceprint signal;

[0087] Perform electromagnetic interference filtering processing on the power frequency phase-synchronized original voiceprint signal based on the simulated interference feature waveform to obtain the voiceprint signal with electromagnetic interference filtered out.

[0088] In an embodiment of the present invention, exemplarily, after the server has obtained a standard voiceprint signal with power frequency phase synchronization (such as a synchronization signal obtained by processing a certain power equipment in a substation under ideal operating conditions), it will generate a simulated interference feature waveform through a specific module in the voiceprint feature enhancement model. The model will simulate the electromagnetic interference feature waveforms that may interfere with the original voiceprint signal, that is, the simulated interference feature waveforms, based on a large amount of previous power equipment operation data and the analysis of common electromagnetic interference situations, combined with the characteristics of the current standard voiceprint signal with power frequency phase synchronization. Then, the server uses the generated simulated interference feature waveform to perform electromagnetic interference filtering processing on the original voiceprint signal with power frequency phase synchronization through the voiceprint feature enhancement model. At this time, the model will carefully compare the simulated interference feature waveform with the original voiceprint signal with power frequency phase synchronization. Once it is found that there are parts in the original voiceprint signal that are similar to the simulated interference feature waveform, it is determined that these parts are very likely caused by electromagnetic interference, and then professional algorithms and processing mechanisms are used to remove these suspected electromagnetic interference parts from the original voiceprint signal with power frequency phase synchronization. After such processing, the electromagnetic interference components in the original voiceprint signal are effectively filtered, thereby obtaining a voiceprint signal with electromagnetic interference filtered, which can more purely reflect the voiceprint characteristics actually emitted by the power equipment and provide a more accurate basis for further analyzing the equipment operating state in the future.

[0089] In an embodiment of the present invention, the electromagnetic interference filtering process of the original voiceprint signal with power frequency phase synchronization based on the simulated interference feature waveform to obtain the voiceprint signal with electromagnetic interference filtered can be implemented through the following examples.

[0090] Perform electromagnetic interference filtering processing on the original voiceprint signal with power frequency phase synchronization based on the simulated interference feature waveform to obtain a preprocessed voiceprint signal of the original voiceprint signal;

[0091] Generate an interference suppression frequency domain matrix for the preprocessed voiceprint signal according to the standard voiceprint signal, the simulated interference feature waveform, the original voiceprint signal, and the preprocessed voiceprint signal;

[0092] Perform frequency domain feature reconstruction processing on the preprocessed voiceprint signal based on the interference suppression frequency domain matrix to obtain the voiceprint signal with electromagnetic interference filtered.

[0093] In an embodiment of the present invention, exemplarily, after receiving the original voiceprint signal with power frequency phase synchronization, the server starts the electromagnetic interference filtering process through the voiceprint feature enhancement model. This model identifies and removes the electromagnetic interference components in the original voiceprint signal based on the generated simulation interference feature waveforms. The model performs spectral analysis on the original voiceprint signal with power frequency phase synchronization, decomposing it into components of different frequencies. At the same time, similar spectral analysis is also performed on the simulation interference feature waveforms. Then, by comparing the characteristic parameters such as amplitude and phase of these two at each frequency component, it is determined which frequency components in the original voiceprint signal are highly similar to the simulation interference feature waveforms, and these similar parts are probably caused by electromagnetic interference. The model uses a specific filtering algorithm to filter out these frequency components determined to be electromagnetic interference from the original voiceprint signal, thereby obtaining a preprocessed voiceprint signal of the original voiceprint signal that is relatively purer. This preprocessed voiceprint signal has initially removed the influence of some obvious electromagnetic interference and is closer to the voiceprint characteristics actually emitted by the power equipment. After the server obtains the standard voiceprint signal, the simulation interference feature waveforms, the original voiceprint signal, and the just-generated preprocessed voiceprint signal, it starts to generate an interference suppression frequency domain matrix for the preprocessed voiceprint signal. The server first performs a comprehensive frequency domain analysis on the standard voiceprint signal, the original voiceprint signal, and the preprocessed voiceprint signal respectively, extracting key frequency domain characteristic parameters such as amplitude and phase at different frequency points. At the same time, in-depth analysis is also performed on the simulation interference feature waveforms to clarify their interference feature manifestations at each frequency point. Then, by comparing the differences and correlations of these several signals in the frequency domain and comprehensively considering the changes in their frequency domain characteristics, using complex mathematical algorithms and the rules built into the model, the parameter values for how to adjust the preprocessed voiceprint signal at each frequency point to further suppress interference are calculated. These parameter values are arranged in the order of frequency, forming the interference suppression frequency domain matrix for the preprocessed voiceprint signal. The server uses the generated interference suppression frequency domain matrix to perform frequency domain feature reconstruction processing on the preprocessed voiceprint signal through the voiceprint feature enhancement model. The model will accurately adjust the frequency domain characteristics such as amplitude and phase of the preprocessed voiceprint signal at each frequency point according to the parameter values in the interference suppression frequency domain matrix. For example, if the parameter at a certain frequency point in the matrix indicates that the amplitude should be reduced, the model will use the corresponding algorithm to reduce the amplitude of the preprocessed voiceprint signal at that frequency point. By adjusting the preprocessed voiceprint signal at each frequency point one by one according to the requirements of the interference suppression frequency domain matrix, its frequency domain characteristics are reconstructed and optimized, further removing the remaining electromagnetic interference influence, and finally obtaining a voiceprint signal with electromagnetic interference filtered out. This voiceprint signal with electromagnetic interference filtered out can more accurately reflect the voiceprint characteristics actually emitted by the power equipment, providing a more reliable basis for accurately judging the operating state of the equipment in the future.

[0094] In an embodiment of the present invention, the power frequency interference elimination process is performed on the voiceprint signal with electromagnetic interference filtered based on the voiceprint feature enhancement model to obtain the voiceprint signal with power frequency noise reduced for the original voiceprint signal, which can be implemented through the following examples.

[0095] Perform a power frequency interference elimination process on the voiceprint signal with electromagnetic interference filtered in the power frequency band dimension based on the voiceprint feature enhancement model to obtain the suppression feature signal of the voiceprint signal with electromagnetic interference filtered in the power frequency band dimension;

[0096] Perform a power frequency interference elimination process on the voiceprint signal with electromagnetic interference filtered in the transient feature dimension based on the voiceprint feature enhancement model to obtain the suppression feature signal of the voiceprint signal with electromagnetic interference filtered in the transient feature dimension;

[0097] Perform an integration operation on the suppression feature signal in the power frequency band dimension and the suppression feature signal in the transient feature dimension to obtain the voiceprint signal with power frequency noise reduced.

[0098] In an embodiment of the present invention, for example, after the server obtains the voiceprint signal with electromagnetic interference filtered, the power frequency interference elimination process is carried out by focusing on the power frequency band dimension through the voiceprint feature enhancement model. The model will perform a spectrum analysis on the voiceprint signal with electromagnetic interference filtered to accurately locate the frequency components related to the power frequency band. For example, for the frequency band corresponding to the operating frequency of common power equipment, the model will elaborate on the characteristics such as the amplitude and phase of the voiceprint signal within this frequency band. Then, based on the built-in power frequency interference elimination algorithm, it will identify the abnormal fluctuations or feature deviations caused by power frequency interference. By adjusting the parameters such as the amplitude and phase of these abnormal parts, the influence of power frequency interference is reduced, so as to obtain the suppression feature signal of the voiceprint signal with electromagnetic interference filtered in the power frequency band dimension, making it better reflect the true voiceprint characteristics of the equipment in this frequency band. At the same time, the voiceprint feature enhancement model also performs a power frequency interference elimination process on the voiceprint signal with electromagnetic interference filtered in the transient feature dimension. In the transient feature dimension, the model will analyze the change characteristics of the voiceprint signal in a short period of time, such as instantaneous amplitude jumps and frequency mutations. For these transient features, the model also relies on relevant algorithms to detect the abnormal transient manifestations that may be caused by power frequency interference. Through targeted processing, such as smoothing the unreasonable fluctuations of the instantaneous amplitude, the suppression feature signal of the voiceprint signal with electromagnetic interference filtered in the transient feature dimension is obtained, further purifying the performance of the voiceprint signal in the transient aspect. Through the above two-dimensional processing by the server, it lays a foundation for further integration processing later to obtain the voiceprint signal with power frequency noise reduced, so as to more accurately present the actual voiceprint characteristics of the power equipment and facilitate the accurate judgment of the equipment operating state.

[0099] In an embodiment of the present invention, the acoustic fingerprint signal after electromagnetic interference filtering is subjected to power frequency interference cancellation processing in the power frequency band dimension based on the acoustic fingerprint feature enhancement model, and a suppression feature signal of the acoustic fingerprint signal after electromagnetic interference filtering in the power frequency band dimension can be obtained by the following examples.

[0100] The acoustic fingerprint signal after electromagnetic interference filtering is decomposed from the transient feature dimension power frequency harmonics to the power frequency band dimension to obtain the harmonic component data of the acoustic fingerprint signal after electromagnetic interference filtering in the power frequency band dimension; the harmonic component data of the power frequency band dimension includes the amplitude spectrum component and the phase spectrum component obtained by decomposing the power frequency harmonics of the acoustic fingerprint signal after electromagnetic interference filtering to the power frequency band dimension;

[0101] The amplitude spectrum component is subjected to adaptive parameter estimation optimization processing to obtain an amplitude spectrum component after adaptive parameter estimation optimization, and the phase spectrum component is subjected to adaptive parameter estimation optimization processing to obtain a phase spectrum component after adaptive parameter estimation optimization;

[0102] According to the amplitude spectrum component after adaptive parameter estimation optimization and the phase spectrum component after adaptive parameter estimation optimization, the suppression feature signal in the power frequency band dimension is determined.

[0103] In an embodiment of the present invention, exemplarily, after receiving the voiceprint signal with electromagnetic interference filtered, the server starts to process the signal through the voiceprint feature enhancement model. First, an operation of decomposing it from the transient feature dimension power frequency harmonics to the power frequency band dimension is carried out. The model uses a specific harmonic decomposition algorithm to conduct a detailed analysis of the voiceprint signal with electromagnetic interference filtered in the frequency domain. It will, according to the law of power frequency harmonics, gradually disassemble and map the complex frequency components of the voiceprint signal in the transient feature dimension to the power frequency band dimension. These components detail the amplitude and phase characteristics of the voiceprint signal at each frequency point in the power frequency band dimension. Next, the server uses the voiceprint feature enhancement model to perform adaptive parameter estimation and optimization processing on the obtained amplitude spectrum component and phase spectrum component respectively. The model will use an adaptive algorithm to dynamically adjust the parameters according to a large amount of accumulated voiceprint data of power equipment and the characteristics of the current voiceprint signal. For the amplitude spectrum component, it will analyze whether there are abnormal fluctuations or deviations from the normal range in the amplitude at each frequency point, and then by adjusting relevant parameters, such as the gain coefficient, etc., make the amplitude more conform to the characteristics that should be in this frequency band during the normal operation of the equipment, so as to obtain the amplitude spectrum component after adaptive parameter estimation and optimization. Similarly, for the phase spectrum component, the model will also check whether its phase change is reasonable, and optimize the accuracy of the phase through a similar adaptive adjustment mechanism to obtain the phase spectrum component after adaptive parameter estimation and optimization. Finally, the server determines the suppression feature signal in the power frequency band dimension based on the amplitude spectrum component and phase spectrum component after the optimization process through the voiceprint feature enhancement model. The model will comprehensively consider the cooperative relationship of these two optimized components at each frequency point, and according to the established rules and algorithms, recombine and construct them into a new signal feature, which is the suppression feature signal in the power frequency band dimension. It more accurately reflects the real situation of the voiceprint signal with electromagnetic interference filtered in the power frequency band dimension, effectively eliminates the influence of power frequency interference, and provides a more reliable basis for subsequent further processing and accurate judgment of the operating state of power equipment. It should be noted that for the power frequency harmonic decomposition, specifically, the voiceprint signal with electromagnetic interference filtered can be set as x(t), and it is decomposed to the power frequency band dimension through synchrosqueezing wavelet transform (SST) to satisfy the formula: where f k is the power frequency fundamental wave (50Hz / 60Hz) and its harmonic components (100Hz, 150Hz, etc.), A k (t) and φ k (t) are the amplitude and phase respectively; the power frequency noise residue is suppressed through an adaptive notch filter (transfer function where ρ = 0.95).

[0104] In an embodiment of the present invention, the power frequency interference elimination processing is performed on the voiceprint signal with electromagnetic interference filtered in the transient feature dimension based on the voiceprint feature enhancement model, and the suppression feature signal of the voiceprint signal with electromagnetic interference filtered in the transient feature dimension is obtained, which can be implemented through the following examples.

[0105] Extract the acoustic feature dimension data of the voiceprint signal with electromagnetic interference filtered in the transient feature dimension based on the voiceprint feature enhancement model;

[0106] Generate a transient interference suppression matrix for the acoustic feature dimension data based on the voiceprint feature enhancement model;

[0107] Perform time-domain pulse cancellation on the acoustic feature dimension data based on the transient interference suppression matrix to obtain the suppression feature signal in the transient feature dimension.

[0108] In an embodiment of the present invention, exemplarily, after receiving the voiceprint signal with electromagnetic interference filtered, the server starts to extract the acoustic feature dimension data in the transient feature dimension through the voiceprint feature enhancement model. The model will conduct a detailed analysis of various acoustic features of the voiceprint signal within a short time interval, such as instantaneous amplitude changes, rapid frequency fluctuations, etc. It will accurately capture these transient features according to specific algorithms and rules, and quantify them into a series of data values. These data values constitute the acoustic feature dimension data of the voiceprint signal with electromagnetic interference filtered in the transient feature dimension, just like taking a detailed "snapshot" of the voiceprint signal in the transient aspect and recording its key transient acoustic features. Then, the voiceprint feature enhancement model generates a transient interference suppression matrix based on the extracted acoustic feature dimension data. The model will use specific algorithms and the internal rule mechanism of the model according to the analysis of a large amount of voiceprint data of power equipment and the understanding of the transient features of the current voiceprint signal with electromagnetic interference filtered. It will comprehensively consider the mutual relationship between different features in the acoustic feature dimension data and the possible influence of power frequency interference, so as to generate a transient interference suppression matrix specifically for these data. The internal parameters of this matrix are set to accurately respond to and suppress the transient interference in the acoustic feature dimension data. Finally, the server uses the generated transient interference suppression matrix to perform time-domain pulse cancellation operation on the acoustic feature dimension data through the voiceprint feature enhancement model. The model will compare and calculate each time-domain pulse in the acoustic feature dimension data with the corresponding parameters in the transient interference suppression matrix. When a pulse that meets the cancellation condition is found, the cancellation process is performed according to the rules set by the matrix. After such a series of cancellation operations, the suppression feature signal in the transient feature dimension is finally obtained, making the voiceprint signal with electromagnetic interference filtered more pure in the transient feature dimension and able to more accurately reflect the actual voiceprint features emitted by the power equipment. It should be noted that the transient interference suppression matrix M is defined as the time-frequency domain energy ratio threshold matrix, and the element M i,j satisfies the formula: where is the energy of the fundamental wave component within the time window t j and σ is the standard deviation threshold statistically calculated from historical data.

[0109] In an embodiment of the present invention, performing an integration operation on the suppression feature signal in the power frequency band dimension and the suppression feature signal in the transient feature dimension to obtain the voiceprint signal with reduced power frequency noise can be implemented through the following examples.

[0110] Generating a first voiceprint feature reconstruction coefficient of the suppression feature signal in the power frequency band dimension and a second voiceprint feature reconstruction coefficient of the suppression feature signal in the transient feature dimension based on the voiceprint feature enhancement model;

[0111] The suppressed characteristic signal in the power frequency band dimension and the suppressed characteristic signal in the transient characteristic dimension are synthesized in feature space based on the first voiceprint feature reconstruction coefficient and the second voiceprint feature reconstruction coefficient to obtain the voiceprint signal with reduced power frequency noise.

[0112] In an embodiment of the present invention, exemplarily, after obtaining the suppression characteristic signal in the power frequency band dimension and the suppression characteristic signal in the transient characteristic dimension, the server generates the corresponding reconstruction coefficient through the voiceprint feature enhancement model. For the suppression characteristic signal in the power frequency band dimension, the model will use a specific algorithm to perform in-depth analysis based on its own characteristics, relevant data accumulated in the previous processing process, and preset rules. For example, considering the amplitude and phase changes of the signal at different frequency points, the first voiceprint feature reconstruction coefficient matching it is accurately generated. Similarly, for the suppression characteristic signal in the transient characteristic dimension, the model will also comprehensively analyze its transient characteristics such as instantaneous amplitude jump and frequency mutation, and calculate the second voiceprint feature reconstruction coefficient through a professional algorithm in combination with relevant information. This coefficient is tailored for subsequent synthesis operations. Then, the server uses the generated first voiceprint feature reconstruction coefficient and the second voiceprint feature reconstruction coefficient to carry out feature space synthesis operations through the voiceprint feature enhancement model. The model will organically combine the suppression characteristic signal in the power frequency band dimension and the suppression characteristic signal in the transient characteristic dimension according to the weights and rules specified by these coefficients. During the synthesis process, the contribution of each signal in different feature dimensions will be adjusted according to the coefficients so that they can be perfectly integrated. After such a feature space synthesis operation, the final result is a voiceprint signal with reduced power frequency noise. This signal has significantly improved in removing the influence of power frequency noise, and can more accurately reflect the characteristics of the voiceprint signal actually emitted by the power equipment, providing strong support for the subsequent accurate judgment of the equipment operation status.

[0113] In an embodiment of the present invention, the voiceprint feature reconstruction processing is performed on the voiceprint signal of the reduced power frequency noise based on the voiceprint feature enhancement model to obtain the current power equipment voiceprint of the original voiceprint signal, which can be implemented through the following examples.

[0114] Acquiring a characteristic stability compensation parameter for voiceprint characteristic energy based on the voiceprint characteristic enhancement model;

[0115] The voiceprint signal of the reduced power frequency noise is processed by voiceprint feature reconstruction based on the characteristic stability compensation parameter to obtain the current power equipment voiceprint.

[0116] In an embodiment of the present invention, exemplarily, after receiving the voiceprint signal for reducing power frequency noise, the server obtains the feature stability compensation parameter for the voiceprint feature energy through the voiceprint feature enhancement model. The model is based on the analysis data of the original voiceprint signal and the voiceprint signal after a series of processes (such as electromagnetic interference filtering, power frequency interference elimination, etc.) during the previous processing. It takes into account the energy changes of the voiceprint feature at different stages, such as whether the energy of the voiceprint feature has unreasonable attenuation or enhancement in certain frequency bands due to the processing. Then, according to the built-in algorithm and the experience summary of a large amount of voiceprint data of power equipment, the model accurately calculates the feature stability compensation parameter that can restore the voiceprint feature energy to a state more in line with the actual state of the equipment. This parameter is used to adjust the voiceprint signal to more accurately reflect the actual situation of the power equipment. Next, the server uses the obtained feature stability compensation parameter to perform voiceprint feature reconstruction processing on the voiceprint signal for reducing power frequency noise through the voiceprint feature enhancement model again. The model will perform fine-tuning on the voiceprint signal for reducing power frequency noise in the acoustic feature dimensions such as frequency, amplitude, and phase according to the adjustment method specified by the feature stability compensation parameter. For example, if the feature stability compensation parameter indicates that the voiceprint feature energy needs to be increased in a certain frequency band, the model will use the corresponding algorithm to increase the amplitude of the voiceprint signal in that frequency band, etc., so that the feature in that frequency band is closer to the state that should be in when the equipment is operating normally. After such voiceprint feature reconstruction processing, the current voiceprint of the power equipment of the original voiceprint signal is finally obtained, and this voiceprint can more accurately reflect the current actual operating state of the target power equipment, providing a reliable basis for subsequent operations such as accurately judging the operating state of the equipment.

[0117] In an embodiment of the present invention, the voiceprint feature enhancement model includes a voiceprint feature enhancement component and a lightweight inference component. The voiceprint feature enhancement component is used to optimize the voiceprint feature of the original voiceprint signal, and the lightweight inference component is used to compress the feature extraction unit of the voiceprint feature enhancement component.

[0118] In an embodiment of the present invention, by way of example, the lightweight inference component may adopt Tucker tensor decomposition technology to compress the feature extraction unit: decompose the original 4D convolution kernel (with a size of H×W×C×K) into 1 core tensor (H’×W’×C’×K’) and 4 factor matrices, set the compression ratio to 4:1, and reconstruct the convolution kernel through matrix multiplication during inference, reducing the memory occupancy by 75%, increasing the calculation speed by 3 times, and the accuracy loss <2%. When the server receives the original voiceprint signal (such as from the acquisition of power equipment), it will send it to the voiceprint feature enhancement component in the voiceprint feature enhancement model. This voiceprint feature enhancement component starts to process the original voiceprint signal in detail. First, it will conduct a comprehensive analysis of the original voiceprint signal, including the analysis of acoustic features such as amplitude and phase in different frequency bands. For example, if it is found that the amplitude of the original voiceprint signal is too low in a certain key frequency band, which may affect the subsequent accurate judgment of the equipment status, the voiceprint feature enhancement component will use a specific algorithm to appropriately increase the amplitude of this frequency band to make it more in line with the characteristic performance of the voiceprint of this equipment in this frequency band under normal circumstances. It will also identify and correct some interference features in the original voiceprint signal that may be caused by factors such as the environment, such as removing abnormal fluctuations caused by some slight electromagnetic interference, etc. Through a series of such operations, the voiceprint features of the original voiceprint signal are optimized, making the voiceprint signal more accurately reflect the actual situation of the power equipment. While the voiceprint feature enhancement component is optimizing the original voiceprint signal, the lightweight inference component in the voiceprint feature enhancement model also starts to work. The task of the lightweight inference component is to compress the feature extraction unit of the voiceprint feature enhancement component. It will analyze various parameters, algorithms, and data structures involved in the processing of the feature extraction unit. For example, if it is found that there are some redundant calculation steps or data representation methods that can be simplified in the feature extraction unit, the lightweight inference component will use specific compression techniques to streamline these redundant parts, reducing the consumption of computing resources without affecting the optimization effect of the original voiceprint signal, improving the efficiency of the entire model processing, enabling the voiceprint feature enhancement model to complete the processing of the original voiceprint signal more quickly and efficiently, and finally obtaining a better current voiceprint of the power equipment. It should be noted that the voiceprint feature enhancement model can adopt a dual-channel generative adversarial network (DC-GAN), where the generator consists of a 5-layer convolutional neural network (CNN), the input is the time-frequency spectrogram of the original voiceprint signal, and the output is the denoised voiceprint; the discriminator is a 3-layer long short-term memory network (LSTM) for identifying the differences between real voiceprints and generated voiceprints; during training, device normal / abnormal voiceprint samples (ratio 1:1) are used, the loss function is a weighted combination of the Wasserstein distance and spectral similarity (weight ratio 0.7:0.3), and the number of training iterations ≥104 times until the signal-to-noise ratio (SNR) of the generator output ≥25dB.

[0119] An embodiment of the present invention provides a computer device 100. The computer device 100 includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned non-contact power device voiceprint monitoring method based on deep learning. As Figure 2 shown, Figure 2 is a structural block diagram of the computer device 100 provided by an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To achieve data transmission or interaction, the elements of the memory 111, the processor 112, and the communication unit 113 are electrically connected to each other directly or indirectly. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.

[0120] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best illustrate the principles of the disclosure and its practical applications, to thereby enable those skilled in the art to best utilize the disclosure and to adapt various embodiments with different modifications to suit the particular applications contemplated.

Claims

1. A non-contact power equipment voiceprint monitoring method based on deep learning, characterized in that: include: The voiceprint of the target power equipment is collected in a preset safe area by using a handheld voiceprint collection device to obtain the original voiceprint signal; Inputting the original voiceprint signal into a pre-trained voiceprint feature enhancement model for processing to obtain the current power equipment voiceprint; The final device operating state corresponding to the target power device is determined based on the current power device voiceprint combined with the device state subdivision type library.

2. The method according to claim 1, characterized in that The determining the final device operation state corresponding to the target power device based on the current power device voiceprint combined with the device state subdivision type library includes: Characterize the voiceprint features of the current power equipment to obtain the target voiceprint features; Based on the target voiceprint feature, a plurality of pending device states are obtained from a device state subdivision type library, wherein the device state subdivision type library is a data set of a plurality of subdivided operating states of the same device operating state; Generate a first voiceprint feature parsing instruction based on the voiceprint feature descriptions and key acoustic dimensions respectively corresponding to the multiple pending device states; By means of an acoustic feature analysis unit, based on the voiceprint feature associated parameters included in the first voiceprint feature analysis instruction, the first voiceprint feature analysis instruction is subjected to voiceprint feature analysis to obtain a first feature analysis output; the first feature analysis output includes first voiceprint feature matching strategies for key acoustic dimensions corresponding to the multiple pending device states respectively; The first voiceprint feature matching strategy is executed by the acoustic feature analysis unit to obtain characteristic parameters of the multiple pending device states in key acoustic dimensions from a voiceprint feature database; wherein the key acoustic dimension is a voiceprint characteristic dimension used to identify the multiple pending device states; Based on the current power equipment voiceprint, the voiceprint feature descriptions corresponding to the multiple pending equipment states, and the feature parameters of the multiple pending equipment states in key acoustic dimensions, the final equipment operation state of the current power equipment voiceprint is obtained from the multiple pending equipment states.

3. The method according to claim 2, characterized in that The key acoustic dimensions are obtained in the following way: The second voiceprint feature analysis instruction is used as a control parameter and loaded into the acoustic feature analysis unit for acoustic feature space mapping to obtain an acoustic reference parameter set for identifying different device status subdivision types, wherein the second voiceprint feature analysis instruction includes the following relevant feature dimensions in the voiceprint feature description: the voiceprint feature description of the same device operating state belonging to the device status subdivision type library, and the voiceprint feature description of multiple device status subdivision types in the device status subdivision type library; The key acoustic dimension is obtained from the acoustic reference parameter set.

4. The method according to claim 3, characterized in that The second voiceprint feature analysis instruction is used as a control parameter and loaded into the acoustic feature analysis unit for acoustic feature space mapping to obtain an acoustic reference parameter set for identifying different device status subdivision types, including: By means of the acoustic feature analysis unit, based on the voiceprint feature associated parameters included in the second voiceprint feature analysis instruction, the second voiceprint feature analysis instruction is subjected to voiceprint feature analysis to obtain a second feature analysis output; the second feature analysis output includes a second voiceprint feature matching strategy for identifying different device status subdivision types; The second voiceprint feature matching strategy is executed by the acoustic feature analysis unit to obtain an acoustic reference parameter set for identifying different device status subdivision types.

5. The method according to claim 2, characterized in that: Based on the target voiceprint feature, Get multiple pending device states from the device state subdivision type library, including: Acquire a typical working condition voiceprint sample set corresponding to a plurality of device state subdivision types in the device state subdivision type library, wherein each device state subdivision type corresponds to at least one typical voiceprint sample in the typical working condition voiceprint sample set; Performing voiceprint feature characterization on each typical voiceprint sample in the typical working condition voiceprint sample set to obtain corresponding typical voiceprint features; Performing voiceprint pattern classification on the obtained multiple typical voiceprint features to obtain voiceprint feature clusters corresponding to the multiple device status subdivision types, wherein one voiceprint feature cluster includes at least one typical voiceprint feature, and the at least one typical voiceprint feature includes the voiceprint feature centroid of the voiceprint feature cluster; Taking the voiceprint feature centroid of the voiceprint feature cluster corresponding to each device status subdivision type as the reference voiceprint feature of each device status subdivision type, and determining the first voiceprint matching degree between the target voiceprint feature and the reference voiceprint feature of each device status subdivision type; The multiple first voiceprint matching degrees are sorted in descending order of priority according to the matching degree, and the device status subdivision types corresponding to a preset number of first voiceprint matching degrees with the highest voiceprint matching degrees are screened out as the pending device status.

6. The method according to claim 2, characterized in that Based on the target voiceprint feature, Get multiple pending device states from the device state subdivision type library, including: Performing acoustic feature modeling on the voiceprint feature description of each device status subdivision type in the device status subdivision type library to obtain a corresponding acoustic feature template; Determine the second voiceprint matching degree between the target voiceprint feature and the acoustic feature template of each device status subdivision type; The plurality of second voiceprint matching degrees are sorted in descending order of priority according to the matching degree, and the device status subdivision types corresponding to a preset number of second voiceprint matching degrees with the highest voiceprint matching degrees are selected as the pending device status.

7. The method according to claim 2, characterized in that The method of obtaining the final device operation state of the current power device voiceprint from the multiple pending device states based on the current power device voiceprint, the voiceprint feature descriptions corresponding to the multiple pending device states, and the feature parameters of the multiple pending device states in key acoustic dimensions, includes: According to the preset voiceprint feature fusion rule, a composite voiceprint analysis instruction is generated based on the current power equipment voiceprint, the voiceprint feature descriptions corresponding to the multiple pending equipment states, and the feature parameters of the multiple pending equipment states in key acoustic dimensions; The acoustic parameter description in the composite voiceprint analysis instruction is subjected to acoustic feature modeling by means of the voiceprint feature analysis model to obtain acoustic analysis features, wherein the acoustic parameter description includes: the voiceprint feature descriptions corresponding to the multiple pending device states respectively and the feature parameters of the multiple pending device states in key acoustic dimensions respectively; By using the voiceprint feature analysis model, the voiceprint feature of the current electric power equipment in the composite voiceprint analysis instruction is characterized to obtain a target voiceprint feature; The acoustic analysis feature and the target voiceprint feature are multi-dimensionally integrated through the voiceprint feature analysis model to obtain a composite acoustic feature; Through the voiceprint feature analysis model, multi-stage state decision optimization is performed according to the composite acoustic features to obtain the final equipment operation state of the current power equipment voiceprint, wherein, when determining the state, the confidence assessment result of the state determination process in the current stage is obtained based on the composite acoustic features and the confidence assessment result of the previous stage.

8. The method according to claim 1, characterized in that The inputting the original voiceprint signal into a pre-trained voiceprint feature enhancement model for processing to obtain the current power equipment voiceprint includes: Get the original voiceprint signal; Obtaining a preset standard voiceprint signal; Obtaining a time-frequency domain matching degree of acoustic feature dimensions between the standard voiceprint signal and the original voiceprint signal; Based on the time-frequency domain matching degree, the standard voiceprint signal and the original voiceprint signal are subjected to power frequency phase synchronization processing to obtain a standard voiceprint signal with power frequency phase synchronization and an original voiceprint signal with power frequency phase synchronization; Generate a simulated interference characteristic waveform based on the standard voiceprint signal synchronized with the power frequency phase; Based on the simulated interference characteristic waveform, the original voiceprint signal with power frequency phase synchronization is subjected to electromagnetic interference filtering processing to obtain a preprocessed voiceprint signal of the original voiceprint signal; generating an interference suppression frequency domain matrix for the preprocessed voiceprint signal according to the standard voiceprint signal, the simulated interference characteristic waveform, the original voiceprint signal and the preprocessed voiceprint signal; Performing frequency domain feature reconstruction processing on the preprocessed voiceprint signal based on the interference suppression frequency domain matrix to obtain the electromagnetic interference filtered voiceprint signal; Decomposing the electromagnetic interference filtered voiceprint signal from the transient feature dimension of power frequency harmonics to the power frequency band dimension, and obtaining the power frequency band dimension harmonic component data of the electromagnetic interference filtered voiceprint signal; the power frequency band dimension harmonic component data includes the amplitude spectrum component and the phase spectrum component obtained by decomposing the power frequency harmonics of the electromagnetic interference filtered voiceprint signal to the power frequency band dimension; Performing adaptive parameter estimation and optimization processing on the amplitude spectrum component to obtain an amplitude spectrum component after adaptive parameter estimation and optimization, and performing adaptive parameter estimation and optimization processing on the phase spectrum component to obtain a phase spectrum component after adaptive parameter estimation and optimization; Determine a suppression characteristic signal in the power frequency band dimension according to the amplitude spectrum component optimized by the adaptive parameter estimation and the phase spectrum component optimized by the adaptive parameter estimation; Extracting acoustic feature dimension data of the electromagnetic interference filtered voiceprint signal in a transient feature dimension based on the voiceprint feature enhancement model; Generating a transient interference suppression matrix for the acoustic feature dimension data based on the voiceprint feature enhancement model; Performing time-domain pulse cancellation on the acoustic feature dimension data based on the transient interference suppression matrix to obtain the suppression feature signal in the transient feature dimension; Generate the first voiceprint feature reconstruction coefficient of the suppressed characteristic signal in the power frequency band dimension and the second voiceprint feature reconstruction coefficient of the suppressed characteristic signal in the transient feature dimension based on the voiceprint feature enhancement model; Based on the first voiceprint feature reconstruction coefficient and the second voiceprint feature reconstruction coefficient, the suppressed feature signal in the power frequency band dimension and the suppressed feature signal in the transient feature dimension are synthesized in feature space to obtain a voiceprint signal with reduced power frequency noise; Acquiring a characteristic stability compensation parameter for voiceprint characteristic energy based on the voiceprint characteristic enhancement model; The voiceprint signal of the reduced power frequency noise is processed by voiceprint feature reconstruction based on the characteristic stability compensation parameter to obtain the current power equipment voiceprint.

9. The method according to claim 8, characterized in that The voiceprint feature enhancement model includes a voiceprint feature enhancement component and a lightweight inference component. The voiceprint feature enhancement component is used to perform voiceprint feature tuning on the original voiceprint signal, and the lightweight inference component is used to compress the feature extraction unit of the voiceprint feature enhancement component.

10. A server system, characterized in that: The method comprises a server, wherein the server is used to execute the method described in any one of claims 1 to 9.