High-precision voice recognition and safety monitoring system and method for electric power operation

By using an improved GAN model and dual-stream recognition architecture, combined with microphone arrays and dual-channel noise suppression, the problem of insufficient speech recognition accuracy in power operations was solved, achieving high-precision speech recognition and real-time safety monitoring, reducing the false recognition rate, and ensuring the safety and accuracy of power operations.

CN121331111APending Publication Date: 2026-01-13LIAONING DONGKE ELECTRIC POWER
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511425640.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In power operation scenarios, speech recognition is not accurate enough and has a high false recognition rate, making it difficult to meet the requirements of real-time performance and accuracy, and also posing safety risks.

Method used

An improved GAN model is used in combination with a microphone array and a dual-channel noise suppression model for speech enhancement. A dual-stream generator and a spectral attention discriminator are used to optimize the speech signal. The noise characteristics of the power operation site are combined with acoustic feature stream and text feature stream for recognition and processing. Real-time evaluation and feedback are carried out in conjunction with power safety regulations.

Benefits of technology

It improves the accuracy of voice recognition, enables real-time safety assessment and timely feedback, reduces the false recognition rate, and ensures the safety and accuracy of power operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121331111A_ABST
    Figure CN121331111A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision voice recognition and safety monitoring system and method for electric power operation. The system comprises a voice acquisition module, a recognition processing module, a real-time evaluation module, a feedback output module and a remote interaction monitoring module. The voice acquisition module is used for acquiring a voice instruction of an operator and performing voice enhancement by using an improved GAN model to obtain an enhanced voice instruction; the recognition processing module is used for recognizing the enhanced voice instruction to obtain a text instruction; the real-time evaluation module is used for performing risk analysis on the text instruction according to the safety regulation, the job ticket and the real-time working condition to obtain an analysis result; the feedback output module is used for sending out voice or graphical early warning through an acousto-optic device according to the analysis result so as to guide an operator to take corrective measures; and the remote interaction monitoring module is used for transmitting the enhanced voice instruction, the text instruction, the analysis result and the correction measure to a background for remote safety monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of power safety monitoring, and particularly relates to a high-precision voice recognition and safety monitoring system and method for power operation. BACKGROUND

[0002] In the power operation scene, voice interaction is an important way to realize efficient operation and remote collaboration, but there are a large number of complex noises such as equipment operation noise on the spot, which will seriously interfere with the collection and recognition accuracy of the voice signal. At the same time, the dialect difference of power operation personnel, the specific pronunciation rules of professional terms and the dynamically changing power operation environment further increase the difficulty of voice recognition. The traditional voice recognition method is designed for general scene, and lacks targeted optimization for power operation environment, resulting in low recognition accuracy and high misrecognition rate, which is difficult to meet the real-time and accuracy requirements of power operation on instruction response, and even may cause safety risks due to recognition errors. Therefore, a high-precision voice recognition and safety monitoring system and method for power operation are urgently needed. SUMMARY

[0003] To solve the above technical problems, the application provides a high-precision voice recognition and safety monitoring system and method for power operation, which has the advantages of improving voice recognition accuracy, realizing real-time safety evaluation, providing timely feedback and remote monitoring.

[0004] To achieve the above purpose, the application provides a high-precision voice recognition and safety monitoring system for power operation, which comprises a voice collection module, an identification processing module, a real-time evaluation module, a feedback output module and a remote interactive monitoring module.

[0005] The voice collection module is used to collect the voice instructions of the operation personnel and perform voice enhancement using an improved GAN model to obtain enhanced voice instructions.

[0006] The identification processing module is used to identify the enhanced voice instructions to obtain text instructions.

[0007] The real-time evaluation module is used to analyze the risks of the text instructions according to the safety regulations, operation tickets and real-time working conditions to obtain analysis results.

[0008] The feedback output module is used to issue voice or graphical early warning through a sound and light device according to the analysis results to guide the operation personnel to take corrective measures.

[0009] The remote interactive monitoring module is used to transmit the enhanced voice instructions, text instructions, analysis results and corrective measures to the background for remote safety monitoring.

[0010] Optionally, the voice instruction of the collection worker is collected and voice enhancement is performed by using the improved GAN model to obtain an enhanced voice instruction, which includes:

[0011] The voice instruction is collected in a directional manner using a microphone array, and a dual-channel noise suppression model for a power operation site is used to obtain a voice instruction after noise reduction;

[0012] The improved GAN model is used to repair the lost voice harmonic components in the voice instruction after noise reduction, and the enhanced voice instruction is obtained.

[0013] Optionally, the dual-channel noise suppression model includes an RNN noise suppression module and a frequency band selective filter.

[0014] The RNN noise suppression module is configured to perform spectral analysis on the voice instruction by short-time Fourier transform, convert the voice instruction from time domain to frequency domain, and obtain the spectral features of the voice instruction.

[0015] The frequency band selective filter is configured to dynamically suppress noise in a corresponding frequency band according to the frequency characteristics of power frequency noise.

[0016] Optionally, obtaining the improved GAN model includes:

[0017] The dual-stream generator is used to enhance the noisy voice signal collected at the power operation site to obtain an enhanced voice sample.

[0018] The pre-recorded power operation typical pure voice instruction sample and the enhanced voice sample are discriminated by using a preset spectral attention discriminator.

[0019] The dual-stream generator and the spectral attention discriminator are used for adversarial game optimization, and an improved GAN model is obtained by combining a target function.

[0020] Optionally, the preset spectral attention discriminator is a multi-channel spectral attention discriminator, which is used to discriminate different noise channels in a targeted manner by combining attention weights and dynamically focusing on key damaged frequency bands, wherein the attention weights are:

[0021]

[0022] wherein β k is the attention weight of the kth channel; g k represents the feature map of the kth channel, and the subscript k represents the channel being calculated; GAP is a global average pooling operation, which takes the average value of all spatial positions of the feature map g k to obtain a scalar value; τ is a temperature parameter, ε is a seasonal dynamic factor, and g i is the feature map of the ith channel.

[0023] Optionally, the recognition processing module comprises an acoustic feature flow unit, a text feature flow unit, a gradient inversion unit and a fusion unit.

[0024] The acoustic feature flow unit is configured to extract time-frequency features of the enhanced voice instruction through a convolution layer, and perform dimension compression on the extracted time-frequency features through a pooling layer to obtain acoustic features.

[0025] The text feature flow unit is configured to fuse a preset knowledge graph to obtain initial text features.

[0026] The gradient inversion unit is configured to correct the initial text features by using a gradient inversion model to obtain text features.

[0027] The fusion unit is configured to fuse the acoustic features and the text features to obtain a text instruction.

[0028] Optionally, the method for performing risk analysis on the text instruction to obtain an analysis result is as follows:

[0029] R = w1·(C v1 +C v2 )+w2·T d

[0030] wherein R is a risk coefficient, C v1 is a number of procedure violations, C v2 is a procedure importance degree, T d is a key operation time deviation, and w1 and w2 are weights.

[0031] Optionally, the feedback output module comprises an audible and visual alarm unit, a voice feedback unit and a graphical display unit.

[0032] The audible and visual alarm unit is configured to perform audible and visual alarm by using a combination of a high-brightness red LED stroboscopic device and a buzzer.

[0033] The voice feedback unit is configured to perform voice warning according to audible and visual alarm.

[0034] The graphical display unit is configured to dynamically display safety information.

[0035] The application further provides a high-precision voice recognition and safety monitoring method for electric power operation, comprising:

[0036] Collecting voice instructions of an operator and performing voice enhancement by using an improved GAN model to obtain enhanced voice instructions;

[0037] Recognizing the enhanced voice instructions to obtain text instructions;

[0038] According to the safety regulations, the operation ticket and the real-time working condition, the risk analysis of the text instruction is carried out, and the analysis result is obtained;

[0039] According to the analysis result, the voice or graphical early warning is sent through the acousto-optic device, and the operator is guided to take the correction measures;

[0040] The enhanced voice instruction, the text instruction, the analysis result and the correction measures are transmitted to the background for remote safety monitoring.

[0041] Compared with the prior art, the present application has the following advantages and technical effects:

[0042] The present application realizes real-time recognition of the voice instruction of the power operator through the high-precision voice recognition technology, carries out real-time safety evaluation in combination with the safety rule library, and provides timely feedback and remote monitoring, thereby effectively solving the problems of insufficient voice recognition precision, untimely safety evaluation and lagging feedback in the prior art, and having the advantages of improving the voice recognition precision, realizing real-time safety evaluation, providing timely feedback and remote monitoring. BRIEF DESCRIPTION OF DRAWINGS

[0043] The drawings constituting a part of the present application are used to provide further understanding of the present application, the illustrative embodiments of the present application and the description thereof are used to explain the present application, and do not constitute improper limitation on the present application. In the drawings:

[0044] Figure 1 is a flow chart of a high-precision voice recognition and safety monitoring system for power operation according to an embodiment of the present application;

[0045] Figure 2 is a flow chart of a high-precision voice recognition and safety monitoring method for power operation according to an embodiment of the present application;

[0046] Figure 3 is a training flow chart of an improved GAN model according to an embodiment of the present application;

[0047] Figure 4 is a block diagram of a high-precision voice recognition and safety monitoring system for power operation according to an embodiment of the present application. DETAILED DESCRIPTION

[0048] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0049] It should be noted that the steps shown in the flow chart of the drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flow chart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0050] The embodiment provides a high-precision voice recognition and safety monitoring system for electric power operation. Figure 1 and Figure 4 As shown in the drawings, the system specifically comprises a voice collection module, an identification processing module, a real-time evaluation module, a feedback output module and a remote interactive monitoring module.

[0051] The voice collection module is used for collecting voice instructions of operation personnel and performing voice enhancement on the voice instructions by using an improved GAN model to obtain enhanced voice instructions.

[0052] The identification processing module is used for identifying the enhanced voice instructions to obtain text instructions.

[0053] The real-time evaluation module is used for performing risk analysis on the text instructions according to safety regulations, operation tickets and real-time working conditions to obtain analysis results.

[0054] The feedback output module is used for issuing voice or graphical early warning through a sound and light device according to the analysis results to guide operation personnel to take corrective measures.

[0055] The remote interactive monitoring module is used for transmitting the enhanced voice instructions, the text instructions, the analysis results and the corrective measures to a background for remote safety monitoring.

[0056] Specifically, the system comprises a power operation voice collection module (the voice collection module), a high-precision power operation instruction identification processing module (the identification processing module), a real-time power site operation safety evaluation module (the real-time evaluation module), a power operation safety feedback output module (the feedback output module) and a remote interactive module of power operation site safety data (the remote interactive monitoring module), wherein:

[0057] The power operation voice collection module is used for receiving voice input of a user and converting the voice input into a digital audio signal.

[0058] The high-precision power operation instruction identification processing module is used for performing voice recognition processing on the received audio signal, identifying the voice instruction input by the user and converting the voice instruction into corresponding text information.

[0059] The real-time power site operation safety evaluation module is used for performing real-time evaluation on the safety of the power operation according to the identified text information, generating a safety evaluation result and judging whether there is a safety hazard according to the evaluation result.

[0060] The power operation safety feedback output module is used for giving corresponding voice feedback or graphical feedback in real time according to the safety evaluation result to remind power operation personnel to take necessary safety protection measures.

[0061] The power operation site safety data remote interaction module is used for sending real-time safety evaluation information of a power operation site to a remote monitoring center to realize remote monitoring and data storage.

[0062] Further, the voice instruction of the operation personnel is collected and voice enhancement is performed by using the improved GAN model to obtain the enhanced voice instruction, which includes:

[0063] The voice instruction is collected in a directional manner using a microphone array, and a dual-channel noise suppression model for the power operation site is used to obtain the voice instruction after noise reduction;

[0064] The improved GAN model is used to repair the lost voice harmonic components in the voice instruction after noise reduction to obtain the enhanced voice instruction.

[0065] Specifically, the power operation voice collection module includes a multi-microphone array, which uses beamforming technology to focus on the direction of human voice for collecting omnidirectional sound information to ensure clear and accurate voice input.

[0066] The voice input optimization processing unit uses the improved generative adversarial network (improved GAN model) to optimize the harmonic structure of human voice in combination with dual-channel noise suppression, for noise elimination, echo suppression and other preprocessing of the input voice, and voice enhancement.

[0067] Further, the dual-channel noise suppression model includes an RNN noise suppression module and a band-selective filter.

[0068] The RNN noise suppression module is used to perform spectral analysis on the voice instruction by short-time Fourier transform to convert the voice instruction from time domain to frequency domain and obtain the spectral features of the voice instruction.

[0069] The band-selective filter is used to dynamically suppress the noise in the corresponding frequency band according to the frequency characteristics of the power frequency noise.

[0070] Further, obtaining the improved GAN model includes:

[0071] The dual-flow generator is used to enhance the noisy voice signal collected at the power operation site to obtain enhanced voice samples.

[0072] The pre-recorded typical pure voice instruction samples of power operation and the enhanced voice samples are discriminated by using the preset spectral attention discriminator.

[0073] The dual-flow generator and the spectral attention discriminator are used for adversarial game optimization, and the improved GAN model is obtained in combination with the objective function.

[0074] Further, the preset spectrum attention discriminator is a multi-channel spectrum attention discriminator, and the multi-channel spectrum attention discriminator is used for performing targeted discrimination on different noise channels in combination with attention weights, and dynamically focusing on a key damaged frequency band.

[0075] Further, the recognition processing module comprises an acoustic feature flow unit, a text feature flow unit, a gradient inversion unit and a fusion unit.

[0076] The acoustic feature flow unit is configured to extract time-frequency features of the enhanced voice instruction through a convolution layer, and the extracted time-frequency features are subjected to dimension compression through a pooling layer to obtain acoustic features.

[0077] The text feature flow unit is configured to fuse a preset knowledge graph to obtain initial text features.

[0078] The gradient inversion unit is configured to correct the initial text features by using a gradient inversion model to obtain text features.

[0079] The fusion unit is configured to fuse the acoustic features and the text features to obtain a text instruction.

[0080] Specifically, the enhanced audio in the power operation voice collection module is input into a dual-flow field adaptive recognition architecture optimized for the power operation scene, and a structured operation instruction text conforming to the power operation safety specification is output, and the dual-flow field adaptive recognition architecture comprises:

[0081] A noise-robust acoustic feature flow is specially used for power field device noise and wind and rain sound environments, and robustly extracts key features of the enhanced audio; a power knowledge embedded text feature flow fuses a power professional term knowledge graph to improve the understanding accuracy of special words and instructions; a gradient reversal layer (GRL) with term correction eliminates term misrecognition caused by dialect pronunciation differences and noise; and a safety rule driven fusion and generation module fuses the dual-flow features, analyzes and generates a standard instruction text according to the power safety regulations, and provides a standard input for safety evaluation.

[0082] Further, the feedback output module comprises an audible and visual alarm unit, a voice feedback unit and a graphical display unit.

[0083] The audible and visual alarm unit is configured to perform audible and visual alarm by using a combination device of a high-brightness red LED stroboscopic + a buzzer.

[0084] The voice feedback unit is configured to perform voice warning according to the audible and visual alarm.

[0085] The graphical display unit is configured to dynamically display safety information.

[0086] Specifically, the early warning mechanism includes an audible and visual alarm unit, wherein the audible and visual alarm adopts a combination of high-brightness red LED stroboscopic and a buzzer to ensure clear identification in a complex and noisy power site operation environment; the early warning mechanism is driven and executed in real time by an embedded edge processing module with an NPU;

[0087] The voice feedback unit automatically synthesizes and broadcasts voice alarms and operation instructions in accordance with the power site safety evaluation results, in accordance with the power safety regulations, to immediately remind the power site workers to avoid risks;

[0088] The graphical display unit dynamically displays core safety information such as the current power operation risk assessment level, key safety regulation items, targeted risk control measures, and emergency disposal points through an anti-glare display screen.

[0089] More specifically, the module is used for wireless connection with external devices, can support wireless communication protocols such as Wi-Fi, Bluetooth or 5G, and the collected voice signals are synchronized in real time to a remote monitoring platform through a 5G / LoRa dual-mode communication module.

[0090] The device also includes a data storage module for storing voice data, safety evaluation results, and operation records of power site workers during the power site operation process for subsequent query and analysis.

[0091] The embodiment also provides a high-precision voice recognition and safety monitoring method for power operation, as shown in Figure 2 , specifically comprising:

[0092] Collecting the operation personnel's voice instructions and using the improved GAN model for voice enhancement to obtain enhanced voice instructions;

[0093] Recognizing the enhanced voice instructions to obtain text instructions;

[0094] According to the safety regulations, the work order and the real-time working condition, the text instructions are analyzed to obtain an analysis result;

[0095] According to the analysis result, a voice or graphical early warning is issued through an audible and visual device to guide the operation personnel to take corrective measures;

[0096] The enhanced voice instructions, text instructions, analysis results and corrective measures are transmitted to the background for remote safety monitoring.

[0097] The embodiment will be described in detail below with reference to the accompanying drawings:

[0098] The embodiment specifically includes the following steps:

[0099] Step one, voice input and spectral purification preprocessing of the power operation voice collection module:

[0100] In the power operation site, the environment is often complex, and there are reverberation, background conversation, equipment operation sound and other noise interference. These noises will seriously affect the accuracy of speech recognition. Specifically, in the present embodiment, the speaking user performs language activities such as reading work tickets in an outdoor space. The purpose of the present embodiment is to perform noise reduction processing on the collected speaking user speech data, obtain complete, clear and noise-free speaking user speech records, and perform high-accuracy text recognition based on the speech records.

[0101] The directional microphone array accurately focuses on the direction of human voice through beamforming technology, effectively retains complete high-frequency voice information, and establishes a dynamic mapping relationship between the acoustic noise characteristics of the outdoor environment and the filtering strength through adaptive noise suppression technology. The core lies in flexibly adjusting the filtering parameters according to the real-time monitored acoustic noise intensity, avoiding the loss of voice naturalness and detail information due to excessive filtering.

[0102] Dynamic filtering parameter adjustment: based on the noise intensity mapping, update the filtering parameters according to the following rules:

[0103]

[0104] wherein, α base is the reference filtering coefficient, k is the electromagnetic coupling coefficient, N peak and N avg are the peak and average intensity of the noise, respectively.

[0105] The core link of the spectrum purification preprocessing is the double-channel parallel noise reduction, which aims to selectively separate human voice and power noise, and further improve the quality of the voice signal.

[0106] Channel A: RNN noise suppression model, based on the RNN noise suppression model (RNNoise algorithm) of the power noise sample library. This algorithm performs spectral analysis on the voice signal through short-time Fourier transform (STFT), converts the voice signal from time domain to frequency domain, so that the spectral characteristics of the voice signal can be clearly displayed. The power noise sample library stores a large number of noise samples commonly found in power operation sites, such as power operation equipment noise, etc.

[0107] Channel B: frequency band selective filter, frequency band selective filter is adopted. There are some power frequency noises with specific frequencies in the power operation site, such as device humming and other steady-state noises. These noises often concentrate in specific frequency bands. The frequency band selective filter can dynamically suppress the noise in the corresponding frequency band according to the frequency characteristics of these power frequency noises, while avoiding damage to the high-frequency components of the voice.

[0108] Through the parallel processing of the double channels, the noise can be suppressed from different angles, greatly improving the effect of noise separation, and making the subsequent processed voice signal more pure.

[0109] Step two, improve the GAN voice enhancement optimization in the power operation voice collection module:

[0110] The improved generative adversarial network (GAN) is used to repair the lost harmonic components of the voice in the noise reduction process, and the specific training process of the improved generative adversarial network (GAN) is as shown in Figure 3

[0111] Step 1, network initialization:

[0112] Initialize the weights of the GRU, Kalman filter, dilated convolution and other modules in the G_base low-frequency channel and G_detail high-frequency channel in the double-flow generator;

[0113] Initialize the multi-channel convolution layer of the spectral attention discriminator, including standard convolution, band-stop convolution and high-pass convolution, and the weights of the attention module.

[0114] Step 2, discriminator adversarial training, with spectral attention focusing:

[0115] Sample noisy-pure voice pairs from the power operation data set;

[0116] Optimize the discriminator objective function:

[0117] (1) Input the pre-recorded pure power voice command sample into the discriminator, marked as "real sample";

[0118] (2) Input the enhanced voice sample output by the generator into the discriminator, marked as "generated sample";

[0119] (3) Spectral attention mechanism dynamic activation:

[0120] Channel 1 standard convolution: focus on general background noise discrimination;

[0121] Channel 2 band-stop convolution: specifically suppress 50Hz power frequency harmonic interference;

[0122] Channel 3 high-pass convolution: focus on discharge noise discrimination.

[0123] (4) Update the discriminator parameters through back propagation to improve its discrimination ability for power noise specific frequency band.

[0124] Step 3, generator adversarial training, superimposed physical constraints:

[0125] Fix the discriminator parameters;

[0126] Optimize the generator objective function: ​

[0127] (1) Dual-stream generation process:

[0128] G_base channel: Generate smooth fundamental frequency trajectory using GRU and Kalman filter to eliminate periodic distortion caused by power frequency noise;

[0129] G_detail channel: Repair high-frequency harmonic structure through dilated convolution and residual attention to fill the friction sound cavity caused by discharge pulse.

[0130] (2) Inject physical constraint terms:

[0131] Fundamental frequency stability constraint: Punish abnormal jumps in fundamental frequency trajectory to prevent vocal distortion of safety instructions;

[0132] Friction sound randomness constraint: Force the energy distribution of plosive and fricative sounds to conform to the non-stationary characteristics;

[0133] KL divergence constraint: Ensure that the randomness of the enhanced speech in the 4-8 kHz frequency band is consistent with the real pronunciation.

[0134] (3) Update generator parameters through backpropagation to improve its ability to recover key speech instructions in strong noise environments.

[0135] Step 4, iterative convergence and deployment:

[0136] Repeat steps 2-3 for adversarial game training.

[0137] Termination condition: The generator loss function contains the optimal stable adversarial loss and physical constraint terms; the enhanced speech passes the power speech recognition system test;

[0138] Save the final generator weights and deploy them to the power operation site speech enhancement equipment.

[0139] (1) Dual-stream generator architecture, power noise presents non-uniform damage in time-frequency domain:

[0140] Low frequency, power frequency noise overwhelms the fundamental frequency F0, causing periodic distortion of speech; high frequency, random discharge pulse destroys the harmonic structure, erasing the friction sound characteristics;

[0141] Traditional single generator is difficult to handle both types of damage, resulting in "fundamental frequency jitter" and "high-frequency cavity" in reconstructed speech.

[0142] 1) G_base structure:

[0143] Input: Low-frequency slice of noise spectrum, focusing on processing power frequency noise interference;

[0144] Use gated recurrent unit (GRU) to model the continuity of the fundamental frequency:

[0145] GRU temporal encoder: 6-layer GRU network, hidden units 256, learn the fundamental frequency dynamic variation law:

[0146] L t = GRU(x t , L t-1 ; X rec )

[0147] where L t represents the output state of the current time step t, which is the output of the GRU unit at the current time; x t represents the input feature of the current time step t; L t-1 represents the output state of the previous time step t-1, which is used to pass historical information; X rec is the recursive weight matrix, which suppresses power frequency interference through the gating mechanism.

[0148] Fundamental frequency smoother: adopt Kalman filter to constrain the fundamental frequency variation rate of adjacent frames, eliminate the fundamental frequency jitter caused by power frequency interference.

[0149] Output: smooth fundamental frequency trajectory signal.

[0150] 2) G_detail structure:

[0151] Input: high-frequency slice of noise spectrum, aiming at discharge pulse noise;

[0152] Dilated convolution pyramid: 4 groups of convolution layers, dilated rates are 1 / 2 / 4 / 8, receptive field is expanded to 320ms, dilated convolution stack is used, dilated factor [1,2,4,8] to capture wide frequency context:

[0153]

[0154] where x is the input feature, usually the time-frequency representation of speech or audio; Conv dil=r represents a dilated convolution layer, where r is the dilated rate; y represents the output feature after 4 layers of dilated convolution stack.

[0155] Residual attention module: calculate high-frequency energy weight, strengthen friction sound area energy weight:

[0156]

[0157] where F(t,f) is the energy value at time frame t and frequency point f; σ is an activation function, used to generate attention weight; B(t,f) is the calculated attention weight matrix; y comes from the output feature of the dilated convolution pyramid; out is the final output, which is the weighted feature plus the original feature.

[0158] Output: repaired high-frequency harmonic details, filling high-frequency holes.

[0159] (2) Spectrum attention discriminator, power noise is non-uniformly distributed in time-frequency domain, traditional discriminators cannot focus on key frequency bands, so a multi-channel spectrum attention mechanism is designed:

[0160] Multi-channel noise specific discrimination, as shown in Table 1:

[0161] Table 1

[0162] Channel Convolution configuration Target noise type Channel 1 Standard 3x3 convolution General background noise Channel 2 Band-stop convolution kernel (stop band 50Hz octaves) Power frequency harmonics Channel 3 High-pass convolution kernel (cutoff 2kHz) Arcing noise

[0163] Attention weight calculation: spectrum attention mechanism, dynamically focusing on key damaged frequency bands.

[0164]

[0165] where: β k represents the attention weight of the kth channel, reflecting the importance of the frequency band, the greater the weight, the more critical the channel; g k represents the feature map of the kth channel; GAP is the global average pooling operation, which takes the average value of all spatial positions of the feature map g k , to get a scalar value; τ is the temperature parameter; ε is the seasonal dynamic factor.

[0166] (3) Physical constraint adversarial training, adversarial game optimization goal:

[0167]

[0168] G generator, its input is noisy speech x noise , the goal is to generate enhanced "clean" speech G(x noise ); D discriminator, its goal is to distinguish between real clean speech and generator generated speech; y clean is the label of the real clean speech; x noise is the input noisy speech; expectation, representing the average of a large number of data samples; min G max D minimax game, the goal of the generator G is to minimize this loss, while the goal of the discriminator D is to maximize this loss.

[0169] In this specific implementation, facing noisy environments, this embodiment aims to ensure that the processed speech sounds natural and conforms to the laws of sound propagation, while also being clear and easy to understand, especially suitable for the characteristics of human voices in power operation scenarios, thereby improving the accuracy of key command recognition. To this end, this embodiment adds an additional constraint based on the characteristics of human hearing to the physical constraints. This is the biophysical constraint. The physical constraints and biophysical constraints are related as a whole to a part, a goal to an implementation method; they are closely linked and jointly serve the core goal of improving GAN speech enhancement, thereby preventing speech distortion.

[0170] 1. Formant continuity constraint to suppress vowel distortion:

[0171]

[0172] The smaller this value, the smoother and more natural the resonance peak; F0 is the trajectory of the fundamental frequency; θ1 is the overweight parameter that controls the importance of this loss.

[0173] 2. Harmonic time-varying characteristic constraint to maintain the randomness of deaf fricatives:

[0174]

[0175] Loss of harmonic time-varying characteristics; The estimated high-frequency harmonic detail components typically refer to the high-frequency random noise characteristics of clean fricatives (such as / s / , / sh / , / f / ); θ2 is a parameter that controls the overweighting of this loss.

[0176] 3. KL divergence constraint:

[0177]

[0178] For loss based on KL divergence; D KL (p∥q) is the KL divergence, used to measure the difference between two probability distributions p and q. The smaller the value, the closer the two distributions are; p real Energy distribution of the enhanced speech in the 4-8kHz frequency band; The actual energy distribution of the unvoiced fricative in the 4-8kHz frequency band; θ3 is the parameter that controls the overweighting of this loss term.

[0179] Based on the above, a mathematical model of the biophysical constraints is performed:

[0180]

[0181] The second derivative of the fundamental frequency trajectory penalizes anomalous jumps;

[0182] Time difference of harmonic components, forced non-stationary characteristics;

[0183] KL divergence term, D KL (p||q) is the KL divergence used to measure the difference between two probability distributions p and q, and the constraint is that the energy distribution of the clean audio segment (4-8 kHz) conforms to the true random characteristics.

[0184] Total generator loss:

[0185]

[0186] Adversarial loss Make the generated speech as real as possible to "fool" the discriminator D;

[0187] Biophysical constraint loss Ensure that the generated speech conforms to the physical laws of sound and prevents distortion.

[0188] Step three, double-flow field adaptive recognition and processing in the high-precision power operation instruction recognition processing module:

[0189] The enhanced audio is input into the double-flow architecture, and the structured instruction text conforming to the power specification is output. The core is to fuse acoustic features and domain knowledge to eliminate dialect differences.

[0190] Acoustic feature flow, extract time-frequency features of speech signals through convolutional layers, such as linear predictive coding cepstrum coefficients LPCC and Mel frequency cepstrum coefficients MFCC. The extracted time-frequency features are compressed in dimension through the pooling layer. The pooling layer can reduce the dimension of feature data, reduce the amount of calculation, and at the same time, it can also retain key feature information and capture local time-frequency patterns of speech, providing effective acoustic feature support for subsequent recognition processing.

[0191] Text feature flow, embed the power operation terminology knowledge graph model, and convert power professional terms into high-dimensional vectors through the knowledge graph embedding layer. The professional terms in the power field have specific meanings and usage scenarios, and traditional speech recognition models often have difficulties in recognizing these terms. Knowledge graph embedding can convert the semantic information of professional terms into high-dimensional vectors that computers can understand, so that the model can better understand and recognize these professional terms, solving the term recognition problem.

[0192] Since power operation personnel may come from different regions and there are dialect differences, a gradient reversal layer (GRL) is introduced as an adversarial training module to eliminate dialect differences.

[0193] Deployment location: after the fully connected layer of the acoustic feature flow;

[0194] Mechanism: identity pass during forward propagation, gradient negation during backward propagation

[0195]

[0196] where: This is the gradient of the loss function α with respect to the network's underlying features x without GRL; δ is the base scaling factor, which determines the basic strength of gradient reversal; γ is the discount factor.

[0197] Effect: forces the model to ignore regional pronunciation differences, such as the northern and southern dialect variants of "insulator".

[0198] During the backward propagation of the model, the gradient reversal layer reverses the gradient direction. This makes the model during training not pay attention to the regional differences in pronunciation, but more focus on the common features in different dialect pronunciation, realizing the invariance learning of dialect features. In this way, the model can better adapt to the pronunciation characteristics of different dialects, and improve the recognition accuracy of different dialects.

[0199] The dual-flow features are fused through the Transformer encoder. The Transformer encoder establishes the association between acoustic features and text features through multi-head attention mechanism, so that the two kinds of features can complement and enhance each other. The multi-head attention mechanism can simultaneously focus on features information in different positions and dimensions, improving the effect of feature fusion.

[0200] Finally, the structured instruction text is generated by the CRF conditional random field decoder. CRF constrains the legality of the output sequence through a transition probability matrix, for example, it can prohibit the appearance of instruction combinations that violate safety procedures, further eliminating recognition errors caused by dialects, and ensuring that the output instruction text meets the specifications and requirements of power operation.

[0201] Step four, real-time evaluation and early warning of safety semantics by the power field operation safety real-time evaluation module:

[0202] Compliance detection: real-time comparison of voice recognition content with the "Electric Power Safety Work Procedures" and the on-site operation standard operating procedures (SOP) clause library to check whether the operation instructions or descriptions conform to relevant regulations and standards. For example, when working on site, if the voice recognition detects the instruction "operate without verifying the electricity", the system will immediately determine that the instruction does not meet the compliance requirements.

[0203] Procedural semantic flow analysis uses Bi-LSTM (Bidirectional Long Short-Term Memory Network) model to learn and model the logical dependency of typical power operation sequence. In power operation, there is a strict logical order and dependency between operations, for example, when performing line maintenance, power must be turned off, checked, grounded, and then subsequent maintenance operations can be performed. Bi-LSTM model can verify the rationality and coherence of the current instruction in the operation process context according to the historical operation sequence and the current voice instruction or state description. If the current instruction is found to have logical contradiction with the previous operation sequence, an early warning will be given in time.

[0204] Operation sequence dependency based on Bi-LSTM modeling:

[0205] z t =LSTM(x t ,z t-1 ),z′ t =LSTM(x t ,z′ t+1 )

[0206] Logical conflict detection, when "line live detection" appears after "grounding line is connected", conflict flag is triggered.

[0207] The power field safety rule library stores the safety operation procedures, risk pre-control measures, typical violation behavior library and equipment safety operation standards dedicated to the power operation site. These information is an important basis for safety evaluation, and provides rich reference materials for compliance detection and context verification.

[0208] Dangerous assessment algorithm, based on the voice recognition result, combined with the information in the field safety rule library for comparative analysis, to identify and warn the potential dangerous factors in the operation process in time. For example, when the operation personnel mentions "high temperature of the equipment", the dangerous assessment algorithm will combine the equipment safety operation standard to assess the risk that may be caused by this situation, and give an early warning in time.

[0209] Input: voice instruction text + equipment state sensor data;

[0210] Output: define dangerous coefficient:

[0211] R=w1·(C v1 +C v2 )+w2·T d , weight w1=0.6, w2=0.3

[0212] Where: C v1 : number of rule violations; C v2 : rule importance; T d : key operation time deviation.

[0213] Step five, early warning and feedback in power operation safety feedback output module:

[0214] Risk label and grading early warning: According to the degree of violation of voice recognition content, the corresponding risk label is generated, such as "general violation", "serious violation" and "serious violation", and the corresponding sound and light alarm is triggered. Sound and light alarm uses LED stroboscopic and buzzer, different risk levels correspond to different alarm modes and frequencies, so that power operation personnel can quickly identify the severity of risk.

[0215] Voice feedback unit can compare and analyze the results automatically generate voice prompts to prompt the relevant errors of power operation personnel according to the safety evaluation results. For example, during the reading process, if it is detected that the key content read by the power operation personnel has serious deviation or omission from the standard ticket, the voice prompt is triggered immediately to remind the power operation personnel to correct the error in time.

[0216] The sound and light alarm unit is shown in Table 2:

[0217] Table 2

[0218] Risk level Light signal Sound signal Low risk Blue constant Single beep Medium risk Yellow flashing Intermittent beep High risk Red strobe Continuous beep

[0219] After the whole work ticket is read, the system summarizes the comparison results. If there are important parts that have not been read, a prompt is issued. And for the high-risk omissions or errors found in the comparison, it is emphasized. Ensure that the work ticket is completely and accurately repeated orally, which is the core link of safety briefing. Through instant voice feedback, it is mandatory to require the reader to supplement or correct the missing / error information on the spot, to ensure that all participants in the power operation can clearly hear the correct safety requirements.

[0220] The graphical display unit displays the power operation risk level, safety information and matters needing attention on the screen. The graphical display is more intuitive and clear, and the power operation personnel can quickly understand the safety status of the current operation and the problems that need attention through the screen.

[0221] Comparison and analysis results, display the real-time risk level obtained according to the comparison results on the screen in a prominent position, clearly list the places where the voice content is inconsistent with the standard work ticket, especially the omitted safety measures, operation steps or risk prompts. Based on the difference analysis, dynamically generate the matters that need to be paid special attention to or immediately supplemented and confirmed by the operation personnel. And the core content of the current standard work ticket can be displayed at the same time as a reference.

[0222] The graphical display mode provides a clear, intuitive, and persistent visual reference for all power operation personnel. It allows everyone to understand the accuracy and completeness of the current reading content at a glance, highlights the existing safety information gaps and potential risks, guides the next action, and serves as a visual record of the safety disclosure process.

[0223] Step six, communication and data management in the power operation site safety data remote interaction module:

[0224] The site safety data remote interaction module is designed specifically for the complex environment of power operation sites and has strong anti-interference capability, supporting wireless connection with external devices and remote monitoring platforms. The module can flexibly select and stably operate in Wi-Fi, Bluetooth, 5G, or LoRa wireless communication protocols according to the network coverage and real-time requirements of the operation site.

[0225] For the collected key operation voice signals, low-latency and high-reliability transmission is realized through the 5G / LoRa dual-mode communication module, and real-time synchronization to the power safety remote monitoring platform is achieved. 5G communication has the characteristics of high speed and low latency, which can meet the real-time data transmission requirements. LoRa communication has the advantages of long transmission distance and strong anti-interference capability, which is suitable for some remote operation sites with poor network coverage. Through dual-mode communication, key data can be transmitted to the remote monitoring platform in a timely and accurate manner, allowing management personnel to monitor the site operation in real time.

[0226] The device also includes a data storage module with shockproof and wide-temperature-range designed storage, which can reliably store the original voice data of the entire operation process, real-time generated safety evaluation results, and key device operation instruction records of the operation personnel in the complex environment of power operation sites.

[0227] These stored data provide detailed basis for subsequent accident traceability analysis. When a safety accident occurs, management personnel can find out the cause and process of the accident by reviewing relevant data, provide reference for procedure optimization by improving safety procedures and operation standards according to the problems encountered in actual operation, and also serve as a basis for personnel performance evaluation to evaluate the work performance and safety awareness of operation personnel.

[0228] In summary, the power operation high-precision voice recognition and safety evaluation device provided in the embodiment can realize real-time monitoring and safety protection of power operation sites through a series of voice processing, recognition, evaluation, and warning links, greatly improving the safety and standardization of power operation. The device has significant advantages in voice recognition accuracy, noise suppression capability, dialect adaptability, and real-time and accuracy of safety evaluation, and can effectively cope with the complex environment and diverse needs of power operation sites.

[0229] The above merely provides the preferred embodiment of the present application, and the protection scope of the present application is not limited thereto. Any modification or replacement within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A high-precision voice recognition and safety monitoring system for power operations, characterized in that, include: The system includes a voice acquisition module, a recognition and processing module, a real-time evaluation module, a feedback output module, and a remote interactive monitoring module. The voice acquisition module is used to acquire the voice commands of the operators and use the improved GAN model to enhance the voice and obtain the enhanced voice commands. The recognition processing module is used to recognize the enhanced voice command and obtain the text command. The real-time assessment module is used to perform risk analysis on text instructions based on safety procedures, work permits, and real-time operating conditions, and to obtain analysis results. The feedback output module is used to issue voice or graphic warnings through an audio-visual device based on the analysis results, guiding operators to take corrective measures. The remote interactive monitoring module is used to transmit the enhanced voice commands, text commands, analysis results, and corrective measures to the backend for remote security monitoring.

2. The high-precision voice recognition and safety monitoring system for power operations according to claim 1, characterized in that, The system collects voice commands from operators and enhances the voice using an improved GAN model. The enhanced voice commands include: Voice commands are collected directionally using a microphone array and combined with a dual-channel noise suppression model for power operation sites to obtain noise-reduced voice commands. The improved GAN model is used to repair the lost speech harmonic components in the denoised speech command to obtain the enhanced speech command.

3. The high-precision voice recognition and safety monitoring system for power operations according to claim 2, characterized in that, The dual-channel noise suppression model includes: an RNN noise suppression module and a frequency band selective filter; The RNN noise suppression module is used to perform spectral analysis on the voice command through short-time Fourier transform, converting the voice command from the time domain to the frequency domain, and obtaining the spectral characteristics of the voice command. The frequency selective filter is used to dynamically suppress noise in a corresponding frequency band based on the frequency characteristics of power frequency noise.

4. The high-precision voice recognition and safety monitoring system for power operations according to claim 2, characterized in that, Obtaining the improved GAN model includes: A dual-stream generator was used to enhance noisy speech signals collected at power operation sites, and enhanced speech samples were obtained. A preset spectral attention discriminator is used to distinguish between pre-recorded typical clean voice command samples for power operations and the enhanced voice samples; The dual-stream generator and spectral attention discriminator are used for adversarial game optimization, and an improved GAN model is obtained by combining the objective function.

5. A high-precision voice recognition and safety monitoring system for power operations according to claim 4, characterized in that, The preset spectral attention discriminator is a multi-channel spectral attention discriminator. This multi-channel spectral attention discriminator is used to combine attention weights to perform targeted discrimination of different noise channels, dynamically focusing on key damaged frequency bands. The attention weights are: Where, β k The attention weight for the k-th channel; g k This represents the feature map of the k-th channel, where the subscript k indicates the channel whose weight is being calculated; GAP is the global average pooling operation, applied to the feature map g. k The average value is obtained by averaging all spatial locations; τ is the temperature parameter, ε is the seasonal dynamic factor, and g i Let be the feature map of the i-th channel.

6. The high-precision voice recognition and safety monitoring system for power operations according to claim 1, characterized in that, The recognition processing module includes: an acoustic feature flow unit, a text feature flow unit, a gradient inversion unit, and a fusion unit; The acoustic feature flow unit is used to extract time-frequency features of enhanced speech commands through convolutional layers. The extracted time-frequency features are then compressed in dimension by pooling layers to obtain acoustic features. The text feature stream unit is used to fuse a preset knowledge graph to obtain initial text features; The gradient inversion unit is used to correct the initial text features using the gradient inversion model to obtain text features; The fusion unit is used to fuse the acoustic features and text features to obtain text instructions.

7. The high-precision voice recognition and safety monitoring system for power operations according to claim 1, characterized in that, The method for performing risk analysis on the text instructions and obtaining the analysis results is as follows: R=w1·(C v1 +C v2 )+w2·T d Where R is the risk factor, and C v1 C represents the number of times the procedure was violated. v2 According to the importance of the procedure, T d The critical operation time deviation is represented by w1 and w2, which are weights.

8. A high-precision voice recognition and safety monitoring system for power operations according to claim 1, characterized in that, The feedback output module includes: an audible and visual alarm unit, a voice feedback unit, and a graphical display unit; The sound and light alarm unit is used to provide sound and light alarms using a combination of a high-brightness red LED flashing and a buzzer. The voice feedback unit is used to issue voice warnings based on the sound and light alarms; The graphical display unit is used to dynamically display security information.

9. A high-precision voice recognition and safety monitoring method for power operations implemented according to the system described in any one of claims 1-8, characterized in that, include: Collect voice commands from operators and use an improved GAN model to enhance the voice, thereby obtaining enhanced voice commands; The enhanced voice command is recognized to obtain the text command; Based on safety regulations, work permits, and real-time operating conditions, risk analysis is performed on text instructions to obtain analysis results; Based on the analysis results, voice or graphic warnings are issued through audio-visual devices to guide operators to take corrective measures; The enhanced voice commands, text commands, analysis results, and corrective measures are transmitted to the backend for remote security monitoring.

Citation Information

Patent Citations

  • Method, apparatus and equipment for establishing voice enhancement network and computer storage medium

    CN109147810A

  • High-frequency sensitive GAN network for LDCT image denoising

    CN110517198A

  • Speech recognition text enhancement system fused with multi-modal semantic invariance

    CN113270086A

  • Electric power safety anti-error management and control method and system based on intelligent voice recognition

    CN118155616A

  • Adaptive traffic domain service voice generation method and system based on thinking chain fine-tuning large model

    CN120126484A