Substation remote intelligent patrol method and system based on voice assistant

By adopting a remote intelligent inspection system based on voice assistant in the operation and maintenance management of substations, the problems of insufficient equipment monitoring, high cost and low security in the traditional operation and maintenance mode are solved, and efficient and safe operation and maintenance management are achieved.

CN120048248APending Publication Date: 2025-05-27BEIJING SGITG ACCENTURE INFORMATION TECH CO LTD +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510069490.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The traditional substation operation and maintenance management model has problems such as insufficient equipment monitoring intensity, high labor costs, low monitoring efficiency, insufficient safety and high learning costs.

Method used

The remote intelligent patrol method and system based on voice assistant is adopted to improve security and recognition accuracy through voice recognition technology and voiceprint detection technology, and to achieve automated patrol, alarm processing and remote control through intelligent module collaboration.

Benefits of technology

It improves the accuracy of speech recognition and the security of the system, enhances the convenience of human-computer interaction and operation and maintenance efficiency, and reduces labor costs and learning costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048248A_ABST
    Figure CN120048248A_ABST
Patent Text Reader

Abstract

According to the transformer substation remote intelligent patrol method and system based on the voice assistant provided by the invention, a plurality of problems in operation and maintenance of a traditional transformer substation are solved by improving a voice recognition technology, enhancing system safety and optimizing module cooperation, and the method and the system have remarkable advantages in the aspects of improving operation and maintenance efficiency and safety; and reliable technical support is provided for intelligent operation and maintenance of the power grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power operation and maintenance, and in particular to a remote intelligent inspection method and system for a substation based on a voice assistant. Background Art

[0002] With the rapid development of the national economy, the scale of substation equipment of the State Grid Corporation has been continuously expanding. The contradiction between the current operation and maintenance management mode and the rapid growth of equipment has become increasingly prominent, and there are problems such as insufficient support and guarantee capabilities for equipment monitoring intensity. The operation and maintenance management fineness is insufficient. The State Grid Corporation implements centralized substation monitoring, optimizes the operation and maintenance monitoring mode of substations, and accelerates the construction of substation centralized control stations.

[0003] The current operation and maintenance management mode is facing contradictions with the rapid growth of equipment. Traditional operation and maintenance management modes often rely on manual monitoring and management. This method is feasible when the number of equipment is small, but as the number of equipment surges, manual monitoring becomes increasingly difficult to cope with, especially in the power system. The safe and stable operation of the power system is of crucial importance, and any mistake may lead to serious consequences.

[0004] At present, many power enterprises have begun to implement centralized substation monitoring to optimize the operation and maintenance monitoring mode of substations and accelerate the construction of substation centralized control stations. However, these traditional monitoring systems still have the following deficiencies: 1) High labor cost: Although the centralized monitoring method is adopted, a large amount of manual participation is still required, increasing the labor cost. 2) Low monitoring efficiency: Traditional monitoring systems often rely on manual monitoring and cannot discover and handle equipment failures in real time and efficiently. 3) Insufficient security: Due to the lack of effective identity verification means, traditional monitoring systems are vulnerable to unauthorized access and there are security risks. 4) High learning cost: Employees need to spend a lot of time learning and mastering complex monitoring systems, increasing the training cost.

[0005] In order to improve the monitoring efficiency and reduce the labor cost, multi-dimensional intelligent management tools have been added. Some power enterprises have begun to explore the use of artificial intelligence technology to assist monitoring. Some have tried to use voice assistants for equipment control, but these technologies still have the following limitations: 1) Low voice recognition accuracy: Existing voice recognition technologies have a low recognition rate in noisy environments, especially in industrial environments, where equipment noise and background noise will affect the recognition effect. 2) Security issues: Lack of effective identity verification mechanisms, and anyone can control equipment through voice assistants, posing security risks. 3) Inconvenient interaction: Existing voice assistant functions are limited and cannot support complex task issuance and result query. 4) Insufficient scalability: Existing systems usually can only perform preset tasks, lack the ability of intelligent learning, and are difficult to adapt to diverse operation and maintenance needs. Summary of the Invention

[0006] To solve the above problems, the present invention proposes a substation remote intelligent inspection method and system based on a voice assistant. This method and system can improve security through voiceprint detection technology, and improve the accuracy of speech recognition and the convenience of human-computer interaction through a voice assistant model.

[0007] A substation remote intelligent inspection method based on a voice assistant proposed by the present invention includes:

[0008] Step 1: Receive user voice input through a voice assistant module, perform speech recognition on the input voice, and transmit the recognized voice command to an intelligent control module;

[0009] Step 2: The intelligent control module receives and parses the voice command, determines the type of the voice command, and executes the corresponding step; if it is an inspection task command, execute Step 3; if it is an alarm information query command, execute Step 4; if it is a dashboard data query command, execute Step 5; if it is a remote control command, execute Step 6;

[0010] Step 3: Execute the inspection task and return an inspection report, and then go to Step 7 after completion;

[0011] Step 4: Monitor the equipment status, trigger an alarm when reaching the threshold or abnormal, return the alarm information, match and push a solution, and then go to Step 7 after completion;

[0012] Step 5: Return the visualized key performance indicators and dashboard data, and then go to Step 7 after completion;

[0013] Step 6: Perform voiceprint recognition verification. If the verification is passed, perform the remote control operation; otherwise, reject the remote control request, and then go to Step 7 after completion;

[0014] Step 7: Convert the system response into voice and output it to the user.

[0015] Preferably, in Step 1, receiving voice input through a voice assistant module, performing speech recognition on the input voice, and transmitting the recognized voice command to an intelligent control module specifically includes:

[0016] Step 1-1: Collect user voice input through a microphone;

[0017] Step 1-2: Perform speech recognition on the voice input to form a text command;

[0018] Step 1-3: Transmit the text command to the intelligent control module through an API interface.

[0019] Preferably, in Step 1-2, performing speech recognition on the voice input to form a text command specifically includes:

[0020] Step 1-2-1: Preprocess the voice input:

[0021] 1-2-1-1: Perform noise reduction and echo cancellation on the input voice to form an audio signal;

[0022] 1-2-1-2: Dynamically adjust the noise reduction intensity according to the environmental changes of the voice input to ensure the clarity of the sound signal;

[0023] Step 1-2-2: Extract features from the processed audio signal:

[0024] 1-2-2-1: Use the Mel Frequency Cepstral Coefficient (MFCC) optimization algorithm to extract the voice features in the audio signal;

[0025] 1-2-2-2: Based on the deep neural network (DNN) enhancement model, further extract the voice features to form multi-dimensional voice features;

[0026] Step 1-2-3: Convert the audio signal from voice to text to form a text instruction:

[0027] 1-2-3-1: Convert the voice feature signal into a text instruction;

[0028] 1-2-3-2: If the effectiveness of the voice input is detected to be low, adopt a secondary confirmation mechanism to prompt the user to re-enter the voice input.

[0029] Preferably, step 4, monitor the device status, trigger an alarm when reaching the threshold or an abnormality occurs, return the alarm information, match and push the solution, specifically includes:

[0030] Step 4-1: Store the historical solutions to form a solution database;

[0031] Step 4-2: Monitor the device status, trigger an alarm when reaching the threshold or an abnormality occurs, and return the alarm information;

[0032] Step 4-3: Retrieve the matching solution in the solution database according to the current event;

[0033] Step 4-4: If there is a matching solution in the system, automatically push the solution to the user; if not, require the user to upload the solution after the event is solved to update the solution database.

[0034] Preferably, step 6, perform voiceprint recognition verification, if the verification is passed, perform the remote control operation, otherwise reject the remote control request, specifically includes:

[0035] Step 6-1: Based on the voiceprint database, compare the input voice with the pre-stored voiceprints to confirm the user's permissions;

[0036] Step 6-2: If the verification is passed, perform the remote control operation according to the instruction and record the operation process; otherwise, execute Step 6-3;

[0037] Step 6-3: Reject the remote control request.

[0038] The present invention also proposes a substation remote intelligent inspection system based on a voice assistant, including:

[0039] A voice assistant module, which is used to receive the user's voice input through the voice assistant module, perform voice recognition on the input voice, and transmit the recognized voice instruction to the intelligent control module;

[0040] An intelligent control module, which is used to receive and parse the voice instruction, judge the type of the voice instruction, and start the corresponding function module; if it is an inspection task instruction, start the intelligent inspection module, if it is an alarm information query instruction, start the intelligent alarm module, if it is a dashboard data query instruction, start the dashboard module, and if it is a remote control instruction, start the remote control module;

[0041] An intelligent inspection module, which is used to execute the inspection task and return the inspection report, and after completion, execute the response output module;

[0042] An intelligent alarm module, which is used to monitor the device status, trigger an alarm when the threshold is reached or an abnormality occurs, return the alarm information, match and push the solution, and after completion, execute the response output module;

[0043] A dashboard module, which is used to return the visual key performance indicators and dashboard data, and after completion, execute the response output module;

[0044] A remote control module, which is used to perform voiceprint recognition verification. If the verification is passed, perform the remote control operation; otherwise, reject the remote control request, and after completion, execute the response output module;

[0045] A response output module, which is used to convert the system response into voice and output it to the user.

[0046] Preferably, the voice assistant module, which is used to receive the user's voice input through the voice assistant module, perform voice recognition on the input voice, and transmit the recognized voice instruction to the intelligent control module, specifically includes:

[0047] A voice input unit, which is used to collect the user's voice input through a microphone;

[0048] A voice recognition unit, which is used to perform voice recognition on the voice input to form a text instruction;

[0049] An instruction transfer unit for transferring the text instruction to the intelligent control module through the API interface. Preferably,

[0050] The remote control module is used to perform voiceprint recognition verification. If the verification is passed, remote control operations are performed; otherwise, the remote control request is rejected. Specifically, it includes:

[0051] A comparison unit for comparing the input voice with the pre-stored voiceprint based on the voiceprint database to confirm user permissions;

[0052] An execution unit for performing remote control operations through voice instructions and recording the operation process if the verification is passed; otherwise, the rejection unit is executed;

[0053] A rejection unit for rejecting the remote control request.

[0054] In summary, the present invention provides a method and system for remote intelligent inspection of a substation based on a voice assistant. Through voice recognition, intelligent module collaboration, and multi-dimensional data processing technologies, many pain points in the traditional operation and maintenance mode are solved. Its beneficial effects are specifically reflected in the following aspects:

[0055] First, improve the accuracy of voice recognition and adapt to voice input in complex environments. The present invention improves the voice recognition technology, significantly improving the recognition accuracy of the voice assistant in complex noise environments. In particular, a feature extraction method based on the Mel Frequency Cepstral Coefficient optimization algorithm and a deep neural network enhancement model is proposed, realizing effective adaptation to low signal-to-noise ratio environments and being able to dynamically adjust feature weights according to signal quality, further enhancing the expression ability of multi-dimensional features such as speech rate and intonation, thereby significantly improving the accuracy of voice recognition. These technical optimizations ensure that the system can accurately recognize user voice instructions even in environments with strong industrial background noise.

[0056] Second, strengthen the security and operability traceability of the system. The present invention proposes a security mechanism based on voiceprint recognition verification to ensure that only legitimate users can execute remote control instructions. Even when the user identity verification fails, the system can reject the operation request, avoiding the execution of unauthorized instructions. This voiceprint-based identity verification method effectively prevents malicious access or incorrect operations. In addition, the present invention also records each operation and result of the user in the remote control module through operation logs, forming a traceable operation chain, further enhancing the security and reliability of the system.

[0057] In addition, the human-computer interaction experience in complex scenarios is guaranteed. The present invention introduces a secondary confirmation mechanism. When the effectiveness of the voice input is detected to be low, the system will prompt the user to re-enter the command through the voice assistant module. This design can effectively avoid task failures caused by input quality problems and improve the fluency and fault tolerance of human-computer interaction.

[0058] Finally, the operation and maintenance efficiency is improved, and the system scalability and adaptability are enhanced. Through the task allocation of the intelligent control module and the cooperation of each functional module, the present invention realizes the full-process automation of substation management. Through modular design, it supports the flexible expansion and upgrade of each functional module. The system can customize functions according to the actual needs of different substations, such as adding a monitoring module for specific equipment or expanding data analysis functions. At the same time, the background platform (such as the 3D data center), combined with real-time monitoring data and historical operation records, can dynamically optimize the alarm threshold and update the solution database to provide reliable support for equipment operation and maintenance in different scenarios.

[0059] In summary, through the improvement of voice recognition technology, the strengthening of system security, and the optimization of module cooperation, the present invention solves many problems in traditional substation operation and maintenance, has significant advantages in improving operation and maintenance efficiency and security, and provides reliable technical support for the intelligent operation and maintenance of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, but do not constitute an improper limitation of the present invention. In the drawings:

[0061] Figure 1 is the method flow chart of the present invention.

[0062] Figure 2 is the system framework diagram of the present invention.

[0063] Figure 3 is the technical architecture diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0064] The present invention will be described in detail below in conjunction with the drawings and specific embodiments, in which the illustrative embodiments and descriptions are only used to explain the present invention, but not to limit the present invention.

[0065] As Figure 1 shown, a substation remote intelligent inspection method based on a voice assistant proposed by the present invention includes:

[0066] Step 1: Receive voice input through the voice assistant module, perform voice recognition on the input voice, and transmit the recognized voice command to the intelligent control module.

[0067] Specifically, first, in step 1-1, the user inputs voice commands through the voice assistant module on a portable device or a fixed device, such as "Execute equipment inspection", "Query the current status", etc. Subsequently, in step 1-2, after the voice assistant module collects the user's voice signal through the microphone, it performs speech recognition on the voice input to form a text command. Finally, in step 1-3, for subsequent processing convenience, the text command will be standardized and then transmitted to the intelligent control module through the API interface.

[0068] Step 2: The intelligent control module receives and parses the voice command, determines the type of the voice command, and executes the corresponding steps; if it is an inspection task command, execute step 3; if it is an alarm information query command, execute step 4; if it is a dashboard data query command, execute step 5; if it is a remote control command, execute step 6;

[0069] Specifically, after receiving the text command transmitted by the voice assistant module, the intelligent control module can call the natural language processing (NLP) module to parse it. Through keyword extraction and semantic analysis techniques, the module can identify the core content in the command (such as "inspection", "alarm", or "remote control"), and execute corresponding tasks according to the command type. For example, if it is an inspection task command, start the automatic inspection process; if it is an alarm information query, call the monitoring function; if it is a dashboard data query, return the corresponding visualization data; if it is a remote control command, verify the user's permission and execute the control operation.

[0070] In step 3, when performing the inspection task, the system calls tools such as remote monitoring devices or drones to complete the equipment inspection according to the preset inspection route or task list. During the inspection process, use cameras and sensors to collect equipment status data (such as temperature, vibration, and operating parameters, etc.), and analyze the data in real time. When an abnormality is found, an alarm prompt will be automatically generated. After the task is completed, record the status and results of the inspection task, generate an inspection report, the content of which includes equipment status data, abnormality records, and processing suggestions. The report will be stored in the database for query.

[0071] In the alarm information query task in step 4, the system will start the equipment status monitoring function, collect the operating data of the equipment in real time and compare it with the preset threshold. If some parameters exceed the threshold or an abnormality occurs, the system will trigger the alarm mechanism, generate an alarm information including the name of the abnormal equipment, the abnormal value, and the reason. At the same time, according to the current event, match and push the optimal processing solution to the user from the solution database that pre-stores historical solutions. If there is no corresponding solution, after the event is resolved, require the user to upload the solution to update the solution database.

[0072] Step 5: For the dashboard data query task, the system calls key performance indicators (KPIs) and dashboard data from the substation monitoring system, such as real-time operating parameters (e.g., voltage, current, etc.), historical trend data, and fault records. The system visualizes this data and presents it to the user in the form of charts and summaries for quick understanding and decision-making.

[0073] Step 6: In the remote control instruction task, the system first performs voiceprint recognition verification by extracting the user's voiceprint features and comparing them with the authorized voiceprints stored in the database to confirm the user's identity. If the verification passes, the system performs corresponding remote control operations, such as device startup or shutdown, operating parameter adjustment, and emergency power-off and other safety operations; if the verification fails, the request is rejected and the user is prompted.

[0074] Step 7: After completing the above tasks, the system converts the results into voice and feeds them back to the user. According to the different results of the inspection report, alarm information, dashboard data, or remote control operations, the system generates corresponding text responses and converts them into voice output through text-to-speech (TTS) technology. According to the type of response, such as alarm information, inspection report, etc., the tone, speed, and volume of the feedback are optimized to improve the user experience. At the same time, the user can also view the detailed information on the interface to ensure that the user can clearly understand the system feedback content.

[0075] The present invention constructs an intelligent remote inspection method for substations by combining voice recognition, natural language processing, adaptive data analysis, and identity authentication. In a complex environment, this method can dynamically adjust the processing strategy and provide real-time feedback on task results, significantly improving the efficiency and safety of substation management, and having broad application prospects and practical value.

[0076] In the method of the present invention, steps 1-2: Performing voice recognition on the voice input is one of the core steps, and its specific implementation method is as follows:

[0077] Step 1-2-1: Preprocess the voice input. Preprocessing of the voice input is a key step to improve the signal quality. First, the system collects the user's voice input through a high-sensitivity microphone. For the collected voice signal, the system uses adaptive noise reduction and echo cancellation technologies for optimization. The noise reduction algorithm can automatically detect the background noise level in the environment and filter out low-frequency noise and interference signals by dynamically adjusting the noise reduction intensity to ensure the clarity of the voice signal. At the same time, the echo cancellation technology can effectively remove audio distortion caused by device echo or reflected sound, further enhancing the quality of the voice.

[0078] The system will also adjust the processing strategy according to the dynamic changes of the environment. For example, when it detects that the user is in a high-noise environment, it automatically enhances the noise reduction parameters to improve the signal clarity; while in a quiet environment, the noise reduction intensity will be appropriately weakened to maintain the natural tone of the voice.

[0079] Step 1-2-2: Extract features from the processed audio signal. After completing the speech preprocessing, the system extracts features from the optimized audio signal to capture the key information contained in the speech, ensuring that the subsequent speech-to-text conversion process can be accurately completed. The feature extraction process specifically includes the following sub-steps:

[0080] 1-2-2-1: Use the Mel Frequency Cepstral Coefficient (MFCC) optimization algorithm to extract the speech features in the audio signal; preferably use the improved MFCC optimization algorithm to extract the spectral features from the audio signal. The MFCC algorithm can effectively simulate the way the human ear perceives sound and generate feature vectors that can reflect the core characteristics of speech. The improved optimization algorithm adds an environmental noise adaptive compensation mechanism during the traditional MFCC extraction process to enhance the stability of feature extraction in a high-noise environment, thereby improving the accuracy of subsequent recognition.

[0081] Mel Frequency Cepstral Coefficient (MFCC) is a commonly used audio feature extraction method, widely used in speech recognition. However, the traditional MFCC algorithm is highly dependent on background noise and speech quality, and often exhibits inaccurate feature extraction or loss of key information in the case of fast speech, accents, or poor speech quality. Therefore, how to introduce a stronger noise adaptability and sensitivity adjustment mechanism on the basis of traditional MFCC to enhance the feature extraction accuracy is the key to improving the speech recognition accuracy. For this reason, the present invention specifically designs a Mel Frequency Cepstral Coefficient (MFCC) optimization algorithm in this step:

[0082] Step 1-2-2-1-1, apply formula (1) to process the audio signal according to the noise level of the input audio, and dynamically adjust the distribution of the frequency axis;

[0083] Adaptive Mel-Frequency=f Mel (SNR)·MFCC baseline (1)

[0084] Wherein, Adaptive Mel-Frequency is the frequency feature after adaptive adjustment, which is the frequency coefficient dynamically adjusted according to the noise level of the audio signal, reflecting the enhancement of the key information in the signal. By adjusting the distribution of the frequency axis, the effective features of the speech signal can be better retained under different noise conditions, thereby optimizing the feature extraction process, especially in a low signal-to-noise ratio (SNR) environment.

[0085] The signal-to-noise ratio SNR represents the ratio of the signal strength to the noise strength. Generally, a higher SNR means a clearer speech signal and less noise interference. Conversely, a low SNR indicates greater noise interference. f Mel (SNR) is a scaling factor for adjusting the frequency axis based on the signal-to-noise ratio (SNR). This factor dynamically adjusts the Mel frequency distribution by analyzing the noise level of the input audio signal. Specifically, at low SNR, the distribution of the frequency axis is adjusted to a way that can better suppress the influence of noise to highlight the key information of the speech; while in a high SNR environment, the distribution of the frequency axis remains relatively standard to emphasize the clarity of the speech signal. For example, at high SNR, f Mel (SNR) takes a value close to 1, indicating that the frequency distribution remains basically unchanged. At low SNR, f Mel (SNR) takes a value greater than 1 or less than 1, indicating corresponding compression or expansion of the frequency axis distribution to enhance the speech features and suppress the noise components.

[0086] MFCC baseline is the baseline Mel-frequency cepstral coefficients, which are the original Mel-frequency cepstral coefficients extracted from the input audio signal without noise adjustment. The baseline MFCC is a spectral feature representation of the audio signal and is usually used in tasks such as speech recognition. In this formula, the baseline MFCC will be multiplied by the frequency scaling factor f Mel (SNR) obtained by adjusting through the noise level to generate the adjusted frequency features adapted to different noise environments.

[0087] This step optimizes the quality of the speech signal by introducing a dynamic adjustment factor f Mel (SNR) based on the noise level to weight the baseline Mel-frequency cepstral coefficients, thereby optimizing the features of the audio signal according to different noise conditions. In a low SNR environment, adjusting the frequency distribution helps reduce the interference of noise on speech recognition; in a high SNR environment, it ensures that the clear features of the speech signal are better retained and optimizes the accuracy of speech recognition.

[0088] Step 1-2-2-1-2, apply formula (2) to optimize the quality of the speech signal; through the noise detection and echo suppression mechanism, dynamically adjust the pre-weighting coefficient α of MFCC to minimize the interference of noise on recognition during the feature extraction process;

[0089] Enhanced MFCC = α·MFCC original +(1-α)·Noise Reduction Component(2)

[0090] Among them, Enhanced MFCC is the optimized Mel Frequency Cepstral Coefficient, which combines the original MFCC features and the noise reduction component, aiming to improve the accuracy of feature extraction, especially in noisy environments. MFCC original is the original Mel Frequency Cepstral Coefficient, which is the result of traditional feature extraction and is used to represent the spectral characteristics in the audio signal. It is extracted from the original audio signal for subsequent speech recognition processing. The Noise Reduction Component is the noise reduction information calculated through the noise detection and echo suppression mechanism. It is used to weaken the interference of environmental noise on feature extraction and usually includes removing background noise, echo, and other unwanted sound components. By introducing the noise reduction component, the system can minimize the impact of noise and thus improve the quality of MFCC.

[0091] α is the pre-weighting coefficient, whose value range is between [0, 1]. It can be automatically adjusted according to the noise level to control the relative importance of the original MFCC and the noise reduction component. When α = 1, it means using only the original MFCC features without introducing the noise reduction component. When α = 0, it means relying entirely on the noise reduction component without using the original MFCC features. When α is between 0 and 1, it means performing weighted fusion between the original MFCC features and the noise reduction component. The value of α is automatically adjusted by the system according to the current noise environment or signal-to-noise ratio (SNR). For example, when the noise environment is relatively clean, α will approach 1; while in a noisy environment, α will approach 0, relying more on the noise reduction component.

[0092] In this step during the audio signal processing, the relationship between the original MFCC features and the noise reduction component is weighted by the dynamically adjusted pre-weighting coefficient α to obtain an optimized Enhanced MFCC. This method can adjust the quality of feature extraction according to the actual noise situation and improve the stability of the speech recognition system in different environments.

[0093] Step 1-2-2-1-3, apply formula (3), and process the MFCC output through an adaptive filter to further eliminate the interference of low-frequency noise, enhance the speech features in the high-frequency part, and more precisely retain the key parts of the speech signal.

[0094] Adaptive Filtered MFCC = MFCC enhanced ·Filter Coefficient(f)(3)

[0095] Among them, Adaptive Filtered MFCC is the Mel Frequency Cepstral Coefficients (MFCC) processed by an adaptive filter. This result is obtained by further processing the MFCC features that have already been enhanced by the filter, thereby enhancing the speech features in the high-frequency part and weakening the interference of low-frequency noise, and optimizing the quality of the speech signal. MFCC enhanced is the enhanced Mel Frequency Cepstral Coefficients, which is an optimized version of the original MFCC features. Some noise components have been removed through the preprocessing steps, improving the quality of the speech signal. Filter Coefficient (f) is the dynamic filtering coefficient for different frequencies f, which is adjusted in real time according to the noise distribution and characteristics of the audio signal. The dynamic filtering coefficient is a frequency-related coefficient that dynamically adjusts the frequency spectrum of the audio signal. The filtering coefficient f is adjusted in real time according to the noise distribution and signal characteristics of the audio signal, such as noise level, signal-to-noise ratio, and spectral characteristics of the speech, with the aim of suppressing noise, especially low-frequency noise, in different frequency ranges while enhancing the speech features in the high-frequency part. This coefficient can be optimized according to the intensity of the ambient noise, the clarity of the signal, and the frequency distribution of the noise. Specifically, when the value of f is larger, the filter suppresses the low-frequency signal more strongly and enhances the speech signal in the high-frequency part. For example, in a noisy environment, the filtering coefficient will automatically increase, thereby enhancing the speech features in the high-frequency part and reducing the influence of low-frequency noise. When the value of f is smaller or close to zero, the filter suppresses the low-frequency signal weakly and mainly retains the features in the low-frequency part, which is suitable for the situation where the low-frequency part of the signal is more important.

[0096] This step further optimizes the MFCC features by introducing the dynamic filtering coefficient f, removing low-frequency noise and enhancing the speech features in the high-frequency part. The filtering coefficient f is adjusted in real time according to the spectrum and noise conditions of the input signal, so that in different noise environments, the key information of the speech signal can be more accurately retained, ensuring that the filter can optimize the signal quality at each moment.

[0097] In summary, in step 1-2-2-1, the MFCC optimization algorithm is used to extract the speech features in the audio signal, and the robustness and accuracy of the speech recognition system are improved by adding an adaptive mechanism. Each feature will automatically adjust its weight according to its importance and signal quality under different environmental conditions, so that the model can select the most effective features in a complex noise environment, ultimately improving the accuracy of speech recognition.

[0098] 1-2-2-2: Based on a deep neural network enhancement model, further extract speech features, including speech speed, intonation, pronunciation clarity, etc., to form multi-dimensional speech features. To capture the complex information of speech signals, the present invention further combines a deep neural network (DNN) enhancement model to extract speech features from multiple dimensions such as speech speed, intonation, and pronunciation clarity. The DNN model can perform multi-layer non-linear mapping on the input MFCC features to generate a multi-dimensional feature vector with stronger semantic expressiveness. This can enhance the adaptability to complex speech (such as dialects and strong accents) and provide more accurate modeling for different types of speech signals (such as natural language instructions or fixed-format commands).

[0099] Traditional deep neural networks (DNNs) have strong expressiveness in extracting multi-dimensional features of speech, but traditional DNN models often encounter the "curse of dimensionality" problem during processing, that is, when the input feature space is too large, the training efficiency and recognition accuracy of the model will decrease significantly. In addition, when modeling audio signals, traditional DNN models may ignore the long-term dependencies of the signals, resulting in inaccurate recognition. How to enhance the robustness and long-term dependencies of DNNs for speech features is a key issue in this technology. Therefore, the present invention also specially designs a deep neural network enhancement model in this step to solve the spatial and temporal problems:

[0100] Step 1-2-2-2-1, based on a deep neural network (DNN) model, according to the speech feature adaptive learning mechanism, add an adaptive weight adjustment layer, enabling it to automatically adjust the weight value of each feature according to the quality and feature dimension of the input signal, so as to adaptively select the most effective feature combination in different environments. The following formula is specifically used:

[0101] Adaptive Feature Weighting = Feature i ·Weight i (t) (4)

[0102] Among them, Adaptive Feature Weighting is the feature value after adaptive feature weighting, reflecting the feature importance under the current environment and signal quality. The weighted features can better represent meaningful information, thereby improving the performance of the subsequent speech recognition system. Feature i is the i-th feature extracted, which can be speech speed, pitch, volume, frequency, etc. Weight i(t) is the feature weight dynamically adjusted at time t based on signal quality and model requirements, which reflects the importance the model attaches to this feature at the current moment. The weight is automatically adjusted according to the quality of the audio signal (such as noise level, speech clarity) and the specific requirements of the moment, and usually changes dynamically. For example, in a noisy environment, the system may increase the weight of features related to speech clarity (such as pitch and articulation clarity) and reduce the weight of frequency features related to noise.

[0103] Step 1-2-2-2-2, based on the adaptive weight adjustment layer, combine the convolutional neural network (CNN) and the recurrent neural network (RNN) joint model to construct a deep neural network (DNN) enhanced model;

[0104] Output DNN =CNN(MFCC enhanced )+RNN(CNN Features) (5)

[0105] where Output DNN is the comprehensive speech feature output by the deep neural network (DNN) enhanced model, and is the final output feature obtained by the deep neural network (DNN) enhanced model. This output is the result of combining the audio features processed by the convolutional neural network (CNN) and the recurrent neural network (RNN), and is obtained through the comprehensive analysis of the DNN. This output is used for subsequent speech recognition or feature decision-making, representing the comprehensive feature of the audio signal.

[0106] CNN(MFCC enhanced ) is the local spatial feature extracted by the convolutional neural network (CNN) from the enhanced mel-frequency cepstral coefficients. MFCC enhanced is the enhanced mel-frequency cepstral coefficients, which represent the spectral information of the audio signal, capture the local patterns, frequency characteristics, etc. in the audio. The CNN extracts local spatial features from these enhanced MFCC features, especially those detailed features related to the speech content of the audio signal. For example, it can identify local features such as intonation and stress in speech.

[0107] RNN(CNN Features) is the temporal feature extracted by the recurrent neural network (RNN) from the local spatial features. CNN Features are the features extracted by the CNN from the input audio signal. Usually, these features contain the local spatial information of the audio signal. The RNN is used to model the time series dependence of these local features. For example, it analyzes the speech rate, pronunciation fluency, rhythm of speech, etc. The RNN enhances the understanding and processing ability of the time changes in the audio signal by capturing the dependencies in time.

[0108] The enhanced model has a strong ability to capture the spatial features (such as intonation, stress, etc.) and time series features (such as speech rate, pronunciation fluency, etc.) of speech signals. The local features of the audio signal are extracted through the CNN module, and then the long-term dependence of the audio signal is modeled through the RNN module, and finally comprehensive analysis of speech features is carried out.

[0109] Step 1-2-2-2-3, apply an adaptive decision layer to the output comprehensive speech features, and dynamically select the most relevant features for processing according to multiple dimensions in the output features, such as speech rate, intonation, pronunciation clarity, etc., to reduce the interference of redundant information and improve the recognition accuracy.

[0110]

[0111] Among them, Adaptive Decision Layer is the final speech feature obtained. This value is the weighted sum based on all features and is the feature after adaptive adjustment, reflecting the key part of the speech signal. The adaptive decision layer will dynamically select the most relevant feature combination according to the relative importance of the features and the current environmental requirements, optimize the speech recognition process, and reduce the interference of redundant information. n is the number of feature dimensions, that is, the number of feature dimensions used in the model. For example, multiple features such as speech rate, intonation, and volume are extracted, and n is the number of these features. Weight i is the adaptive weight of the i-th feature, indicating the importance of this feature in the decision-making. This weight is usually dynamically adjusted according to factors such as signal quality and noise level. Feature i is the i-th feature, representing information from different aspects extracted from the audio signal. Each feature may represent different dimensions of the audio signal, such as intonation, speech rate, pronunciation clarity, etc.

[0112] In summary, in step 1-2-2-2, based on the deep neural network enhanced model, further extract speech features to form multi-dimensional speech features, and use the adaptive mechanism in the deep learning model to improve the accuracy of the speech recognition system. Each feature will automatically adjust its weight according to its importance and signal quality under different environmental conditions, so that the model can select the most effective features in a complex noise environment, and finally improve the accuracy of speech recognition.

[0113] 1-2-2-3: Real-time monitor parameters such as the volume, speech rate, and signal-to-noise ratio of the speech input, and adjust the feature extraction strategy according to these dynamic characteristics. For example: when it is detected that the speech rate is fast, automatically enhance the ability to capture features of fast speech. When the volume of the speech signal is too low or the background noise is large, increase the noise compensation intensity in feature extraction to ensure the integrity of the signal.

[0114] Step 1-2-3: After feature extraction is completed, the present invention performs speech-to-text conversion on the processed speech signal to generate text instructions that the system can recognize. This process is completed through the following steps:

[0115] Step 1-2-3-1: Use a well-trained deep learning model (such as a long short-term memory network LSTM or an end-to-end ASR model) to convert the input multi-dimensional speech feature signal into corresponding text instructions. The deep learning model utilizes its powerful modeling ability to accurately decode the semantic and structural information in the speech signal and generate high-quality text instructions.

[0116] Step 1-2-3-2: If the system detects that the validity of the speech input is low (for example, there is a lot of background noise or the user has a strong accent), a secondary confirmation mechanism will be triggered. The system uses the voice assistant module to feedback prompt information to the user, such as "Please confirm your instruction again" or "The environmental noise is large, please re-enter the instruction". In this way, the system can effectively avoid instruction recognition errors caused by input quality problems.

[0117] Step 1-3: After generating the text instructions, the system passes the instructions to the intelligent control module through the API interface for parsing and subsequent processing. The API interface provides an efficient and standardized instruction transmission method to ensure seamless connection between the voice assistant module and the intelligent control module. After receiving the instructions, the intelligent control module can perform specific tasks according to the text content, such as starting a patrol inspection, querying alarm information, or performing remote control operations.

[0118] The above steps achieve efficient processing of the user's voice input by introducing steps such as speech preprocessing, feature extraction, speech-to-text conversion, and instruction transmission. Especially through the combination of dynamically adjusting the noise reduction intensity, adaptive feature extraction strategy, and deep learning model, the system can maintain high-efficiency and accurate recognition ability in various complex environments. At the same time, the introduction of the secondary confirmation mechanism further improves the reliability of instruction input. The present invention provides an efficient and stable technical solution for remote intelligent inspection of substations and has significant practical value.

[0119] The present invention also designs a creative improvement for the voiceprint recognition and verification link, specifically including:

[0120] Step 2-1: Input speech feature extraction. This step describes the detailed process of the system extracting features from the user's input speech signal, providing high-precision and multi-dimensional input data for subsequent voiceprint verification. By combining traditional spectral feature extraction algorithms and dynamic characteristic analysis, this step can comprehensively reflect the global and local characteristics of the user's voiceprint. Specifically, it includes the following sub-steps:

[0121] Step 2-1-1: Extract the basic spectral features of the voice signal to generate feature vectors, representing the overall spectral characteristics of the user's voiceprint. The collected signal can be preprocessed to remove background noise and enhance the clarity of the voice signal.

[0122] In the present invention, it is preferably to use the traditional Mel Frequency Cepstral Coefficient (MFCC) algorithm to extract the basic spectral features. First, the system preliminarily processes the voice signal input by the user to ensure the signal quality. Subsequently, the basic spectral features of the voice signal are extracted through the traditional algorithm, and the specific process is as follows:

[0123] First, perform Short-Time Fourier Transform (STFT) to convert the voice signal from the time domain to the frequency domain. Through frame segmentation, the continuous voice signal is divided into small segments, and each segment is subjected to Fourier transform to generate a spectrogram. Secondly, a set of filters based on the Mel scale are used to weight the spectrogram. The Mel scale simulates the perception characteristics of the human ear for signals of different frequencies. By emphasizing the low-frequency part (related to voice information) and weakening the high-frequency part (related to noise), the key information of the voice signal can be effectively extracted. Finally, logarithmic operation is performed on the filter output to obtain the log power spectrum, and then the feature vectors are extracted through logarithmic operation and Discrete Cosine Transform (DCT). These feature vectors are the compressed representation of the voice signal and contain the main information of the voice spectrum. The finally generated feature vectors can comprehensively reflect the overall spectral characteristics of the user's voiceprint, including core elements such as the frequency distribution of pronunciation, intonation, and speech rate, providing basic data for voiceprint comparison.

[0124] Step 2-1-2: On the basis of the basic spectral features, to further improve the system's ability to capture voiceprint details, the system introduces a micro-vibration characteristic extraction unit. By analyzing the dynamic changes of the high-frequency and low-frequency components of the voice signal, dynamic features that can reflect the local characteristics of the voiceprint are generated.

[0125] The system separates the input voice signal into frequency bands, and extracts the high-frequency signal (usually reflecting voice detail characteristics, such as vocal cord vibration) and the low-frequency signal (usually reflecting the overall contour of the voice) respectively. The high-frequency signal F high (t) has a frequency range above 1 kHz and contains fine characteristics of the voice, such as intonation changes and vibration details. The low-frequency signal F low (t) has a frequency range below 1 kHz and contains the basic structural information of the voice, such as the voice contour.

[0126] Regarding the dynamic changes of the high-frequency and low-frequency features, extract the local feature differences of the voice signal. The specific formula is:

[0127]

[0128] Among them, F high (t) and F low(t) are the high - frequency and low - frequency signal intensities, respectively, and t 1 ,t 2 is the time interval, which is used to capture the dynamic change characteristics of voiceprint details. This formula calculates the differential integral of the high - frequency signal and the low - frequency signal, extracts the dynamic characteristics of the user's voiceprint, and reflects the specificity of the voiceprint in the local area.

[0129] Based on the dynamic change calculation, it is also possible to further conduct a joint analysis by combining the two dimensions of time and frequency. By using the comparison of short - time energy change and spectral energy distribution, the complex dynamic characteristics of the speech signal are captured.

[0130] Preferably, the MFCC algorithm is used to obtain the spectral characteristics of the speech signal. The MFCC algorithm provides the overall spectral characteristics of the speech signal, while the micro - vibration characteristic extraction unit enhances the expression ability of local dynamic changes, ensuring the all - round coverage of voiceprint characteristics.

[0131] Through steps 2 - 1 - 1 and 2 - 1 - 2, the system extracts the overall spectral characteristics and local dynamic characteristics of the speech signal respectively, forming a multi - dimensional feature representation. These feature vectors not only contain the core spectral information of the user's voiceprint but also reflect the uniqueness of its pronunciation details, effectively improving the accuracy and robustness of voiceprint recognition, providing high - quality input data for the subsequent verification steps. It also achieves the effect of strong anti - noise ability. Through the dynamic analysis of high - and low - frequency signals and time - frequency joint processing, the system can extract stable voiceprint characteristics under complex background noise, enhancing the robustness of voiceprint verification. Moreover, it has a strong ability to distinguish fake voiceprints. The micro - vibration characteristics capture the fine changes in the user's speech, which are difficult to simulate by forged voiceprints, thus significantly improving the system's resistance to fake voiceprint attacks.

[0132] Step 2 - 2: Voiceprint comparison and fake voiceprint detection. Aiming at the security and accuracy problems in the voiceprint verification process, combined with multi - dimensional feature analysis technology, a fake voiceprint detection module is introduced to comprehensively improve the anti - forgery ability and verification accuracy of the system. Through feature matching, fake voiceprint probability calculation, and comprehensive evaluation of verification confidence, the efficient analysis and accurate verification of the input voiceprint are realized, which specifically includes the following sub - steps:

[0133] Step 2 - 2 - 1: Preliminary voiceprint feature matching.

[0134] Voiceprint database retrieval. The system matches the extracted user voiceprint feature vectors with the pre - stored reference voiceprint data in the voiceprint database. The voiceprint database contains multi - dimensional voiceprint templates, including the user's basic spectral characteristics, dynamic change characteristics, and historical verification records, etc., ensuring that the retrieval and matching results can be obtained quickly and efficiently in different scenarios.

[0135] Calculate the matching degree score. During the matching process, a feature similarity calculation method is adopted. By measuring the Euclidean distance and cosine similarity index between the input voiceprint features and the reference voiceprints in the database, the matching degree score S is generated. match . The specific calculation formula is:

[0136]

[0137] where: Input Feature is the input voiceprint feature vector, Database Feature is the voiceprint template of the corresponding user in the database, and S match is the matching degree score, ranging from [0, 1]. The higher the value, the higher the matching degree.

[0138] Step 2-2-2: Pseudo-voiceprint detection. To counter pseudo-voiceprint attacks, such as simulated voiceprint or recording playback attacks, the system designs a pseudo-voiceprint detection module. The feature differences between real voiceprints and pseudo-voiceprints are learned through a generative adversarial network (GAN) model.

[0139] The GAN model consists of a generator and a discriminator. The generator is responsible for generating the feature distribution of pseudo-voiceprints, a feature vector that is similar to the real voiceprint features but has forgery characteristics. The discriminator learns to distinguish the features of real voiceprints and pseudo-voiceprints by continuously classifying the input features.

[0140] Through adversarial training, the model accurately captures the feature distribution of pseudo-voiceprints and discriminates the input voiceprints, calculating the pseudo-voiceprint probability P fake . The specific discrimination function is:

[0141] P fake = f GAN (Input Features) (9)

[0142] where: P fake is the pseudo-voiceprint probability, ranging from [0, 1]. The higher the value, the more obvious the pseudo-voiceprint characteristics. f GAN is the classification function obtained through adversarial network training.

[0143] To improve the accuracy of pseudo-voiceprint detection, the present invention also constructs a multi-layer classification function f based on adversarial training GAN to capture the differences between pseudo-voiceprints and real voiceprints in the high-dimensional feature space. This function combines multi-dimensional feature fusion, time-frequency characteristic weight adjustment, and an adaptive activation mechanism to ensure robustness and efficiency under complex input conditions.

[0144] f GAN (InputFeatures)

[0145] = σ(W 3 ·ReLU(W2 ·tanh(W 1 ·Input Features + b 1 ) + b 2 )

[0146] +b 3 )(10)

[0147] The classification function f designed in the present invention GAN is a three - layer structure. The activation function of the first layer is the hyperbolic tangent function (tanh), which is used to smooth the input features and is suitable for processing noisy input signals; among them,

[0148] the output range is [-1, 1], making the signal more smooth and contrastive. The first - layer structure is a feature pre - processing layer, which maps multi - dimensional input features to a high - dimensional hidden - layer space, emphasizing the differences in spectral details between pseudo - voiceprints and real voiceprints.

[0149] The activation function of the second layer is the rectified linear unit (ReLU), which is used to strengthen positive activation and suppress the influence of invalid features; among them, ReLU(x) = max(0, x). The second layer is a feature fusion layer, which further extracts the time - frequency characteristics of pseudo - voiceprints and adjusts the feature weights to dynamically adapt to different noise environments.

[0150] The activation function of the output layer is the Sigmoid function (σ), which is used to map the final output to the probability of pseudo - voiceprints; among them, the output range is [0, 1]. The output layer is a pseudo - voiceprint discrimination layer, which finally maps the fused features to the pseudo - voiceprint probability P fake .

[0151] In the formula, the input features (InputFeatures) are the multi - dimensional voiceprint feature vectors extracted, including basic spectral features and micro - vibration dynamic characteristic features, denoted as: InputFeatures = [x 1 , x 2 , …, x n , where n is the dimension of the feature vector.

[0152] (W 1 , W 2 , W 3 , b 1 , b 2 , b 3 ) are the weight matrix and bias terms. W 1 , W 2 , W 3 are the weight matrices of the corresponding layers, and b 1 , b 2 , b 3Is the bias term for the corresponding layer.

[0153] W 1 ∈R m×n : The first layer maps from an n-dimensional feature space to an m-dimensional hidden layer space; W 2 ∈R k×m : The second layer maps from the m-dimensional hidden layer space to a k-dimensional hidden layer; W 3 ∈R 1×k : The output layer maps the k-dimensional hidden layer space to the final pseudo voiceprint probability, and R represents a matrix.

[0154] Specifically, when introducing the time-frequency characteristic weight adjustment in the weight matrix W of the first layer: 1

[0155] W 1 [i,j] = α·F freq [i] + (1 - α)·T time [j] (11)

[0156] Where: F freq [i] is the weight of the i-th dimensional frequency characteristic, T time [j] is the weight of the j-th dimensional time characteristic, and α is a dynamic adjustment coefficient, which is updated in real time according to the environmental noise or the quality of the input features.

[0157] The GAN model further combines the time series characteristics to analyze the dynamic change differences between high-frequency and low-frequency features in the speech signal. Pseudo voiceprints usually lack the dynamic features of the high-frequency part, while the high-frequency features of real voiceprints contain micro-vibration detail changes. Therefore, the model can effectively detect the abnormal characteristics of pseudo voiceprints.

[0158] Through the above innovative design, the pseudo voiceprint classification function f GAN can fully learn the differential features between real voiceprints and pseudo voiceprints, especially the differences in spectral details and time series changes. Its output P fake is the pseudo voiceprint probability, which accurately reflects the authenticity of the input voiceprint and provides high-precision technical support for subsequent comprehensive evaluation. And through the combination of multi-layer non-linear activation functions, the learning ability of the model for complex feature distributions is significantly enhanced, especially the ability to capture pseudo voiceprint details. Using the activation function structure combined with ReLU and tanh takes into account the sparsity and computational efficiency of feature extraction, ensuring the rapid convergence of adversarial training.

[0159] Step 2-2-3: Verify the comprehensive evaluation of confidence.

[0160] The system evaluates the final verification confidence S of the input voiceprint by comprehensively considering the matching score S match and the pseudo voiceprint probability P fake . The calculation formula is as follows: final ​​

[0161] S final = w 1 ·S match - w 2 ·P fake (12)

[0162] Where: w 1 is the weight of the matching score, and w 2 is the weight of the pseudo voiceprint probability, indicating its impact on the final confidence. S final is the final verification confidence, ranging from [-1, 1]. The higher the value, the more credible the input voiceprint is.

[0163] To meet the verification requirements of different scenarios, the system dynamically adjusts the weights w 1 and w 2 according to the noise environment, voice input quality, etc. For example, in a low-noise environment, the weight w match of the matching score S 1 is higher; while in a scenario with a higher probability of pseudo voiceprint attacks, the weight w 2 of the pseudo voiceprint probability will increase accordingly.

[0164] By implementing accurate feature comparison and efficient pseudo voiceprint detection during the voiceprint verification process, it provides a strong technical guarantee for the security and reliability of the system.

[0165] Step 2-3: Verification result processing. Perform identity verification and operation permission control according to the final verification confidence S final .

[0166] If S final ≥ T accept (threshold), the verification passes, and the user is allowed to perform remote control operations;

[0167] If S final < T accept , the verification fails, the system rejects the request, and the voice assistant prompts the user with the reason for failure (such as "Verification failed, please re-enter").

[0168] Step 2-4: Operation record and security feedback. To enhance the traceability of system operations, the system records verification results and operation logs, including voiceprint features input by the user, matching scores, pseudo voiceprint probabilities, and operation times, etc. If a pseudo voiceprint attack is detected, the system will automatically trigger a security alarm, record the attack information in the database, and push it to the operation and maintenance personnel.

[0169] It can be seen that in the voiceprint recognition and verification process of the present invention, by introducing micro-vibration characteristics and time-frequency analysis, the detailed changes in voiceprints are captured, enhancing the verification ability of the system in complex scenarios. The generative adversarial network (GAN) is used to detect fake voiceprints, improving the system's resistance to simulated voiceprint attacks. Moreover, the user experience is optimized. By combining the verification confidence and the dynamic feedback mechanism, the user verification process is ensured to be smooth, while reducing the false rejection rate. Through detailed logging and attack warning mechanisms, the traceability is improved, and the system's security response ability in abnormal situations is enhanced. In summary, the method not only effectively improves the security of voiceprint recognition and verification, but also enhances the protection ability against fake voiceprint attacks in complex scenarios, providing more efficient and reliable technical support for the remote intelligent patrol system.

[0170] As Figure 2 shown, the present invention also proposes a substation remote intelligent patrol system based on a voice assistant. Through the cooperation of multiple modules, the system realizes remote monitoring, inspection, data query and control of substation equipment. The following are the specific implementation manners of each system module: including:

[0171] The voice assistant module is used to receive user voice input, complete voice recognition, and transfer the recognition result to the intelligent control module. Its specific functions include:

[0172] The voice input unit. This unit collects the user's voice input through a high-sensitivity microphone, and optimizes the voice signal using an adaptive noise reduction algorithm and echo suppression technology to ensure clear voice signals can still be obtained in complex noise environments. By real-time monitoring the signal quality, the signal acquisition parameters can be dynamically adjusted.

[0173] The voice recognition unit processes the collected voice signal. First, the features of the voice signal (such as frequency, intonation, etc.) are extracted, and then the voice is converted into a text command through a deep learning model. The recognition unit also supports optimization processing for noise or complex accents. For example, a secondary confirmation mechanism is used to prompt the user to re-enter the command to ensure accurate recognition.

[0174] The command transfer unit transfers the text command to the intelligent control module through the API interface. The transfer process uses a standardized protocol to ensure the efficiency and accuracy of command transmission.

[0175] The intelligent control module is the core part of the system, used to parse voice commands and start the corresponding function modules.

[0176] According to the command type, the intelligent control module can start the following sub-modules:

[0177] Intelligent Patrol Module. The intelligent patrol module is responsible for the automated patrol of equipment. After receiving an instruction, the module calls patrol equipment (such as drones or robots) to collect equipment status data (such as temperature, voltage, etc.) along a preset patrol route. The collected data is analyzed in real time. If an anomaly is detected, the system generates a patrol report containing detailed status information, anomaly records, and recommended handling solutions. After the patrol is completed, the report is transmitted to the response output module.

[0178] Intelligent Alarm Module. The intelligent alarm module is used to monitor the equipment status in real time. When the equipment parameters reach the preset threshold or an anomaly occurs, an alarm is triggered and a solution is pushed. The specific functions are as follows: Storage Unit: Stores historical solutions to form a solution database. Monitoring Unit: Monitors the running status of the equipment in real time and determines whether to trigger an alarm based on the threshold. Matching Unit: Retrieves solutions matching the current event in the solution database. If no matching solution is found, the system prompts the user to upload a solution after the event is resolved to update the database. Push Unit: If a matching solution is found, the system automatically pushes the solution to the user and provides operation guidance.

[0179] Dashboard Module. The dashboard module is responsible for returning visual data of the substation's key performance indicators (KPIs). Through a data query interface, the module obtains the real-time operating parameters (such as current, voltage, etc.) and historical trend data of the equipment, and generates an intuitive chart report. Users can receive this data through the response output module to assist in decision-making.

[0180] Remote Control Module. The remote control module allows users to control equipment through voice commands. To ensure security, the module first performs voiceprint recognition verification to confirm user permissions. The specific functions are as follows: Comparison Unit: Compares the user's voice with the pre-stored voiceprint database to verify the user's identity. Execution Unit: After verification, the module performs remote control operations, including starting the equipment, adjusting parameters, etc., and records operation logs for query. Rejection Unit: If the verification fails, the module rejects the remote control request and sends a prompt message to the user through the response output module. All audio operations will retain the audio and the corresponding recognition conversion files for tracking and auditing. All audio operations will retain the audio and the corresponding recognition conversion files for tracking and auditing.

[0181] Response Output Module. The response output module is used to feedback the processing results of the system to the user, specifically including: Voice feedback: Through text-to-speech (TTS) technology, convert text instructions into voice output to ensure that the user can directly hear the response results of the system. Interface display: Display detailed response content on the user terminal interface, such as inspection reports, alarm information, dashboard data, etc., to provide a more intuitive reference. This module can generate corresponding feedback content according to the completion status of different tasks. For example: In the inspection task, output the summary or key information of the inspection report; in the alarm task, push the reason for the abnormality and the solution; in the remote control task, feedback the execution status of the operation or the reason for rejecting the request.

[0182] The remote intelligent inspection system of the substation of the present invention realizes remote monitoring, automatic inspection, intelligent alarm and safety control of the substation through the collaborative work of the voice assistant module, the intelligent control module and each functional module. The adaptive feature extraction, deep learning model and voiceprint recognition technology in the system ensure high efficiency and safety in complex environments. This system greatly improves the management efficiency and safety of the substation and has wide application value.

[0183] As Figure 3 shown, the overall technical architecture of the remote intelligent inspection system of the substation of the present invention is jointly composed of a voice assistant module, an intelligent inspection module, an intelligent alarm module, a dashboard push module, a remote control module and a background support platform, covering the all-round integration of the data center, the power grid business center and the unified video monitoring platform, and realizing remote monitoring, intelligent operation and maintenance and efficient management of the substation.

[0184] The system architecture and module functions include:

[0185] The voice assistant module is responsible for receiving and parsing user voice instructions, extracting the instruction content through voice recognition technology and transmitting it to the intelligent control module.

[0186] The intelligent inspection module supports identity authentication, voiceprint detection and voice instruction parsing, and can start functions such as inspection, alarm, dashboard push or remote control according to different task requirements.

[0187] The intelligent alarm module monitors based on real-time device status data. When the device operation parameters exceed the preset threshold or an abnormality occurs, the system will trigger an alarm and match a suitable processing solution according to the built-in solution database and push it to the user. The alarm module supports threshold dynamic adjustment, log analysis and performance diagnosis to ensure the safe operation of the device.

[0188] The dashboard module is responsible for the visual push of key performance indicators (KPIs), including real-time device status, historical operation data, and fault analysis results. By analyzing performance indicators and trend data, it helps users intuitively grasp the operation of the substation.

[0189] The remote control module provides the function for users to operate devices through voice commands. Its core functions include: Voiceprint recognition verification: Perform remote operations after confirming user permissions;

[0190] The system background consists of a 3D data center, a power grid resource business center, and a unified video surveillance platform, which provide the following supports respectively: The 3D data center integrates inspection data, solution databases, and log records, supporting status monitoring and multi-dimensional data analysis. The power grid resource business center is responsible for the standardized management of device data, providing device alarm thresholds and command control interface supports. The unified video surveillance platform combines video surveillance data to achieve real-time visual monitoring of device status and supports dynamic linkage operations.

[0191] To improve the usability and maintainability of the system, a practical toolset is also introduced, including: Key management tool: Ensure the security of data transmission and user operations; Collaboration management tool: Support multi-user task assignment and progress management; Operation log: Record and analyze user operations and system responses; Simulation verification tool: Used for testing and optimization before deployment.

[0192] In terms of the work process, the substation remote intelligent inspection system takes the voice assistant module as the entry point and realizes the remote monitoring, inspection, alarm, and control of substation equipment through multi-module collaboration and background support.

[0193] Users issue voice commands through the voice assistant module, such as "Perform equipment inspection" or "Query alarm information". The voice assistant module collects the user's voice signal through the microphone, converts the voice into a text command using noise reduction and speech recognition technologies, especially the specially designed Mel Frequency Cepstral Coefficient (MFCC) optimization algorithm and Deep Neural Network (DNN) enhancement model, and at the same time analyzes the user's operation intention. After the analysis is completed, the text command is transmitted to the intelligent control module through the API interface.

[0194] After receiving the instruction, the intelligent control module starts the corresponding function module according to the instruction type. For the inspection task, the system starts the intelligent inspection module, performs automatic inspection through drones or fixed devices, collects equipment status data and generates inspection reports; for the alarm information query task, the system starts the intelligent alarm module, monitors the equipment status in real time, triggers an alarm if an abnormality is found, and matches the best treatment plan from the solution database and pushes it to the user; if the user queries the dashboard data, the dashboard module will extract key performance indicators (KPIs) and generate a visual report on the real-time operation status and trend analysis; for the remote control instruction, the system starts the remote control module, first verifies the user's identity through voiceprint recognition, and then performs operations such as equipment startup, parameter adjustment or emergency power-off, while recording the operation log.

[0195] After the task is completed, each module transfers the result to the response output module. The response output module uses text-to-speech technology to feedback the result to the user, such as "The inspection is completed, and 3 abnormalities are found", and at the same time displays detailed data on the user terminal interface, including inspection reports, alarm reasons, solution plans or operation logs, etc. Finally, all data is stored in the 3D data center for subsequent analysis, query and optimization.

[0196] Through the above process, the system realizes the full closed-loop management of voice instruction reception, task execution and result feedback. Combined with the data storage and analysis functions of the background platform, it provides an efficient and intelligent remote inspection solution for the substation, significantly improving the operation and maintenance efficiency and safety.

[0197] Through the deep combination of voice interaction technology and intelligent analysis technology, this system realizes the full-process intelligent management of substation operation and maintenance, and has the following advantages: through the cooperation of the voice assistant and intelligent modules, the operation process is simplified and the user experience is improved. Using the 3D data center and real-time monitoring data, it provides comprehensive equipment status management and problem diagnosis capabilities. Adopting voiceprint verification and log management mechanisms to ensure the security and traceability of remote control. Supports modular function expansion and can be customized and deployed according to the needs of different substations.

[0198] In summary, the substation remote intelligent inspection system of this embodiment can efficiently complete the inspection, monitoring, alarm and control of the substation through the application of multi-module linkage and intelligent processing technology, significantly improving the operation and maintenance management efficiency of the power grid and having extensive practical application value.

[0199] In summary, the present invention provides a substation remote intelligent inspection method and system based on a voice assistant. Through voice recognition, intelligent module cooperation and multi-dimensional data processing technology, it solves many pain points in the traditional operation and maintenance mode. Its beneficial effects are specifically reflected in the following aspects:

[0200] First, improve the accuracy of speech recognition and adapt to speech input in complex environments. The present invention significantly improves the recognition accuracy of the speech assistant in complex noise environments by improving speech recognition technology. In particular, a feature extraction method based on the Mel Frequency Cepstral Coefficient (MFCC) optimization algorithm and a deep neural network enhancement model is proposed, which enables effective adaptation to low signal-to-noise ratio environments and can dynamically adjust feature weights according to signal quality, further enhancing the expression ability of multi-dimensional features such as speech rate and intonation, thus significantly improving the accuracy of speech recognition. These technical optimizations ensure that the system can accurately recognize user speech commands even in environments with strong industrial background noise.

[0201] Second, enhance the security and operational traceability of the system. The present invention proposes a security mechanism based on voiceprint recognition verification to ensure that only legitimate users can execute remote control commands. Even when the user authentication fails, the system can reject the operation request, preventing unauthorized commands from being executed. This voiceprint-based authentication method effectively prevents malicious access or incorrect operations. In addition, the present invention also records every operation and result of the user in the remote control module through operation logs, forming a traceable operation chain, further enhancing the security and reliability of the system.

[0202] In addition, ensure the human-machine interaction experience in complex scenarios. The present invention introduces a secondary confirmation mechanism. When the effectiveness of the speech input is detected to be low, the system will prompt the user to re-enter the command through the speech assistant module. This design can effectively avoid task failures caused by input quality problems and improve the fluency and fault tolerance of human-machine interaction.

[0203] Finally, improve the operation and maintenance efficiency, and enhance the scalability and adaptability of the system. The present invention realizes the full-process automation of substation management through the task allocation of the intelligent control module and the cooperation of each functional module. Through modular design, it supports the flexible expansion and upgrade of each functional module. The system can customize functions according to the actual needs of different substations, such as adding a monitoring module for specific equipment or expanding data analysis functions. At the same time, the background platform (such as a 3D data center), combined with real-time monitoring data and historical operation records, can dynamically optimize alarm thresholds and update the solution database, providing reliable support for equipment operation and maintenance in different scenarios.

[0204] In summary, the present invention solves many problems in traditional substation operation and maintenance by improving speech recognition technology, strengthening system security, and optimizing module cooperation, and has significant advantages in improving operation and maintenance efficiency and security, providing reliable technical support for the intelligent operation and maintenance of the power grid.

[0205] The above are only the preferred embodiments of the present invention. Therefore, all equivalent changes or modifications made according to the structures, features and principles described in the scope of the present invention patent application are included in the scope of the present invention patent application.

Claims

1. A remote intelligent inspection method for a substation based on a voice assistant, comprising: Step 1: receiving user voice input through the voice assistant module, performing voice recognition on the input voice, and transmitting the recognized voice command to the intelligent control module; Step 2, the intelligent control module receives and analyzes the voice command, determines the type of the voice command, and executes the corresponding step; if it is an inspection task command, execute step 3; if it is an alarm information query command, execute step 4; if it is a dashboard data query command, execute step 5; if it is a remote control command, execute step 6; Step 3: Perform the inspection task and return the inspection report. After completion, go to step 7; Step 4: Monitor the device status, trigger an alarm when the threshold is reached or an abnormality occurs, return the alarm information, match and push the solution, and go to step 7 after completion; Step 5: Return the visualized key performance indicators and dashboard data. Go to step 7 after completion. Step 6: Perform voiceprint recognition verification. If the verification is successful, perform the remote control operation. Otherwise, reject the remote control request. After completion, go to step 7. Step 7: Convert the system response into speech output to the user.

2. The method of claim 1, wherein: The step 1, receiving voice input through the voice assistant module, performing voice recognition on the input voice, and transmitting the recognized voice command to the intelligent control module, specifically includes: Step 1-1, collecting user voice input through a microphone; Step 1-2, performing voice recognition on the voice input to form a text instruction; Step 1-3: passing the text instruction to the intelligent control module through the API interface.

3. The method of claim 2, wherein: The step 1-2, performing voice recognition on the voice input to form a text instruction, specifically includes: Step 1-2-1: pre-process the voice input: 1-2-1-1, performing noise reduction and echo cancellation processing on the input voice to form an audio signal; 1-2-1-2. Dynamically adjust the noise reduction intensity according to the changes in the voice input environment to ensure the clarity of the sound signal; Step 1-2-2: Extract features from the processed audio signal: 1-2-2-1. Use Mel-frequency cepstral coefficient (MFCC) optimization algorithm to extract speech features from audio signals; 1-2-2-2. Based on the deep neural network (DNN) enhanced model, further extract speech features to form multi-dimensional speech features; Step 1-2-3: Convert the audio signal from speech to text to form text instructions: 1-2-3-1. Convert voice feature signals into text instructions; 1-2-3-2. If the validity of the voice input is detected to be low, a secondary confirmation mechanism is used to prompt the user to re-enter the voice input.

4. The method of claim 3, wherein: Step 1-2-2-1: Use the Mel Frequency Cepstral Coefficient (MFCC) optimization algorithm to extract speech features from the audio signal, specifically including: Step 1-2-2-1-1, apply formula (1), process the audio signal according to the noise level of the input audio, and dynamically adjust the distribution of the frequency axis. Adaptive Mel-Frequency=f Mel (SNR)·MFCC baseline (1) Among them, Adaptive Mel-Frequency is the frequency feature after adaptive adjustment, f Mel (SNR) is a function that adjusts the frequency axis based on the signal-to-noise ratio (SNR), MFCC baseline is the reference frequency cepstrum coefficient; Step 1-2-2-1-2, apply formula (2) to automatically optimize according to the voice signal quality, MFCC enhanced =α·MFCC original +(1-α)·Noise Reduction Component(2) Among them, MFCC enhanced It is the optimized frequency cepstrum coefficient, MFCC original is the original frequency cepstrum coefficient, Noise Reduction Component is the noise reduction component, α is the pre-weighting coefficient, the value range is [0,1], and α is automatically adjusted according to the noise level; Step 1-2-2-1-3, apply formula (3), process the MFCC optimization output through an adaptive filter, further eliminate the interference of low-frequency noise, and enhance the speech characteristics of the high-frequency part. Adaptive Filtered MFCC=MFCC enhanced ·Filter Coefficient(f)(3) Among them, Adaptive Filtered MFCC is the frequency cepstrum coefficient after being processed by the adaptive filter, and FilterCoefficient(f) is the dynamic filter coefficient for different frequencies f.

5. The method of claim 3, wherein: Step 1-2-2-2: Based on the deep neural network (DNN) enhanced model, further extract speech features to form multi-dimensional speech features, including: Step 1-2-2-2-1, based on the deep neural network (DNN) model, according to the speech feature adaptive learning mechanism, add an adaptive weight adjustment layer; according to the quality and feature dimension of the input signal, automatically adjust the weight value of each feature. The specific algorithm is as shown in formula (4). Adaptive Feature Weighting=Feature i ·Weight i (t) (4) Among them, Adaptive Feature Weighting is the feature value after adaptive feature weighting, Feature i is the i-th feature extracted, Weight i (t) is the feature weight dynamically adjusted based on signal quality and model requirements at time t; Step 1-2-2-2-2, based on the adaptive weight adjustment layer, combined with the convolutional neural network (CNN) and recurrent neural network (RNN) joint model, build a deep neural network (DNN) enhancement model. The specific algorithm is as shown in formula (5). Output DNN =CNN(MFCC enhanced )+RNN(CNN Features) (5) Among them, Output DNN It is the speech feature output by the deep neural network (DNN) enhancement model, CNN (MFCC enhanced ) is the local spatial feature extracted by the convolutional neural network (CNN) for the enhanced frequency cepstral coefficients, and RNN (CNN Features) is the temporal feature extracted by the recurrent neural network (RNN) for the local spatial feature; Step 1-2-2-2-3, apply an adaptive decision layer to the speech features output by the enhancement model, and dynamically select the most relevant features for processing based on multiple dimensions in the output features. The specific algorithm is as shown in formula (6). Among them, Adaptive Decision Layer is the final speech feature obtained, n is the number of feature dimensions, Weight i is the adaptive weight of the i-th feature, Feature i is the ith feature.

6. The method according to claim 1, wherein step 4, monitoring the device status, triggering an alarm when a threshold is reached or an abnormality occurs, returning alarm information, matching and pushing a solution, specifically includes: Step 4-1, store historical solutions to form a solution database; Step 4-2: Monitor the device status, trigger an alarm when the threshold is reached or an abnormality occurs, and return the alarm information; Step 4-3: According to the current event, search for matching solutions in the solution database; Step 4-4: If there is a matching solution in the system, the solution is automatically pushed to the user; If not, the user is asked to upload a solution after the incident is resolved, and the solution database is updated.

7. The method according to claim 1, wherein step 6, performing voiceprint recognition verification, if the verification is passed, performing the remote control operation, otherwise rejecting the remote control request, specifically comprises: Step 6-1: Based on the voiceprint database, compare the input voice with the pre-stored voiceprint to confirm the user's authority; Step 6-2: If the verification is passed, the remote control operation is performed according to the instruction and the operation process is recorded; otherwise, step 6-3 is executed; Step 6-3: reject the remote control request.

8. A remote intelligent inspection system for substations based on voice assistant, comprising: A voice assistant module is used to receive user voice input through the voice assistant module, perform voice recognition on the input voice, and transmit the recognized voice command to the intelligent control module; An intelligent control module is used to receive and analyze the voice command, determine the type of the voice command, and start the corresponding function module; if it is an inspection task command, the intelligent inspection module is started; if it is an alarm information query command, the intelligent alarm module is started; if it is a dashboard data query command, the dashboard module is started; if it is a remote control command, the remote control module is started; Intelligent inspection module, used to execute inspection tasks and return inspection reports, and execute response output module after completion; Intelligent alarm module, which is used to monitor the status of equipment, trigger alarms when thresholds are reached or abnormalities occur, return alarm information, match and push solutions, and execute response output modules after completion; The dashboard module is used to return visual key performance indicators and dashboard data, and execute the response output module after completion; A remote control module is used to perform voiceprint recognition verification. If the verification is passed, the remote control operation is performed, otherwise the remote control request is rejected and the response output module is executed after completion; The response output module is used to convert the system response into voice output to the user.

9. The system of claim 4, wherein: The voice assistant module is used to receive user voice input through the voice assistant module, perform voice recognition on the input voice, and transmit the recognized voice command to the intelligent control module, specifically including: A voice input unit, used to collect user voice input through a microphone; A speech recognition unit, used to perform speech recognition on the speech input to form a text instruction; The instruction transmission unit is used to transmit the text instruction to the intelligent control module through the API interface.

10. The system of claim 4, wherein: The remote control module is used to perform voiceprint recognition verification. If the verification is passed, the remote control operation is performed, otherwise the remote control request is rejected, and specifically includes: A comparison unit, used to compare the input voice with the pre-stored voiceprint based on the voiceprint database to confirm the user's authority; An execution unit, used to execute the remote control operation through voice command and record the operation process if the verification is passed; otherwise, the rejection unit is executed; A rejecting unit is used to reject the remote control request.

Citation Information

Patent Citations

  • Multi-target voice enhancement method based on SCNN (Stacked Convolutional Neural Network) and TCNN (Temporal Convolutional Neural Network) joint estimation

    CN110867181A

  • Voice-interaction-based intelligent control method for transformer substation and distribution station

    CN111179928A

  • Text sentiment classification method based on hybrid model

    CN111309909A

  • Power equipment voice recognition method and system in inspection scene

    CN113851116A

  • Inspection system and method based on voice interaction

    CN115019411A