Video and voiceprint dual-mode power equipment monitoring method and device

Through the dual-modal monitoring method of video and voiceprint, combined with high-definition cameras and microphone arrays to collect power equipment data, perform feature extraction and anomaly detection, it solves the shortcomings of single-modal monitoring, achieves high-confidence early warning and data upload, and improves the accuracy and reliability of power equipment monitoring.

CN120808549APending Publication Date: 2025-10-17HUANENG SHAANXI JINGBIAN ELECTRIC POWER CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510846372.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing power equipment monitoring methods rely on a single data source, making it difficult to accurately identify faults in complex environments, resulting in missed reports, false alarms, and low confidence in judgments. They also lack the ability to collaboratively analyze multi-source data and make accurate decisions.

Method used

A dual-modal monitoring method of video and voiceprint is adopted. The video images and voiceprint signals of the power equipment are synchronously collected through high-definition cameras and microphone arrays for preprocessing and feature extraction. The convolutional neural network model is used to combine the equipment appearance and voiceprint features for anomaly detection, realizing dual-modal data cross-validation and high-confidence early warning, and saving the abnormal event data packets and uploading them to the cloud.

Benefits of technology

It improves the accuracy and reliability of power equipment anomaly detection, realizes real-time monitoring and timely early warning, reduces the risk of equipment failure, and improves the safe operation level of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808549A_ABST
    Figure CN120808549A_ABST
Patent Text Reader

Abstract

The invention provides a video and voiceprint dual-mode power equipment monitoring method and device, and the method comprises the steps: synchronously collecting an initial video image and an initial voiceprint signal of the operation of power equipment, carrying out the preprocessing of the initial video image and the initial voiceprint signal, and obtaining a target video image and a target voiceprint signal; performing feature extraction on the target video image and the target voiceprint signal to obtain an equipment appearance feature and an equipment voiceprint feature; based on a convolutional neural network model, performing equipment anomaly detection in combination with the equipment appearance features and the equipment voiceprint features, and if anomaly is detected in a single mode, starting dual-mode data cross validation; if abnormity is detected in the double modes, high-confidence early warning is triggered; storing the complete data packet of the abnormal event, and synchronously uploading to the cloud; and early warning information is sent out, and related personnel are notified through a grading alarm mechanism, so that the equipment fault risk is effectively reduced, and the safe operation level of a power system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of power equipment monitoring, and particularly relates to a video and voiceprint dual-mode power equipment monitoring method and device. BACKGROUND

[0002] In the field of power system operation and maintenance, real-time monitoring of equipment status is of vital importance to ensuring power supply safety and stability. With the continuous expansion of the power grid scale and the increase in equipment complexity, how to effectively identify potential faults and take timely measures has become a core requirement for industry development. The operating status of power equipment is directly related to energy supply and the normal operation of social economy, therefore, it is particularly urgent to study efficient and accurate monitoring technology.

[0003] However, the current mainstream monitoring methods have exposed some deep-seated deficiencies in practical application. The technology relies too much on the analysis of a single data source, making it difficult to cope with changing factors in complex environments, including: Single video monitoring: relies on cameras to capture the appearance and operating status of equipment, but cannot detect internal discharge, insulation aging and other invisible faults. For example, discharge sparks may not be captured due to lighting conditions or obstructions.

[0004] Single voice monitoring: collects sound signals through microphones to analyze abnormalities, but is easily disturbed by environmental noise and cannot directly reflect the appearance status of equipment (such as deformation and damage).

[0005] Therefore, single-mode monitoring is prone to false negatives and false positives, making it difficult to detect early faults in time, and the accurate discrimination ability of signals is weak in a strong interference background, affecting the safety of the power system.

[0006] In addition, existing methods often lack effective integration mechanisms when dealing with multi-source information, resulting in insufficient comprehensive judgment of fault characteristics, missing early hidden dangers or making unnecessary misjudgments, and thus affecting the overall reliability of the system. A deeper challenge is how to achieve collaborative analysis and accurate decision-making of multi-source data. In complex operating environments, video data may be affected by light or obstructions, while sound data is easily overwhelmed by background noise, making it difficult for single-mode feature extraction to achieve ideal results. The problem of information isolation also directly leads to low confidence in fault judgment, especially when facing early weak anomalies, the system often cannot form reliable warning basis, increasing the risk of equipment damage.

[0007] Therefore, how to integrate video and sound data features of two different types to build a monitoring mechanism that can verify each other and make comprehensive decisions has become a key problem to improve the accuracy and reliability of power equipment fault identification. SUMMARY

[0008] The application provides a video and voiceprint dual-mode power equipment monitoring method and device to solve the low accuracy and reliability of power equipment fault identification in the prior art.

[0009] In one aspect, the application provides a video and voiceprint dual-mode power equipment monitoring method, which comprises: Synchronously collecting initial video images and initial voiceprint signals of power equipment operation, and pre-processing the initial video images and initial voiceprint signals to obtain target video images and target voiceprint signals; Extracting features from the target video images and target voiceprint signals to obtain equipment appearance features and equipment voiceprint features; Based on a convolutional neural network model, equipment anomaly detection is performed by combining the equipment appearance features and equipment voiceprint features, if an anomaly is detected by a single mode, cross-validation of dual-mode data is started, and if anomalies are detected by both modes, a high-confidence early warning is triggered; Saving complete data packets of abnormal events and synchronously uploading them to the cloud; Issuing early warning information and notifying relevant personnel through a hierarchical alarm mechanism.

[0010] According to the video and voiceprint dual-mode power equipment monitoring method provided by the application, the process of obtaining target video images and target voiceprint signals comprises: Synchronously collecting operation data of power equipment through a high-definition camera and a microphone array, and recording device appearance video images and sound signals in real time to obtain initial video images and initial voiceprint signals; Segmenting and enhancing the video frames in the initial video images by using image processing technology to obtain target video images; Separating and suppressing the environmental noise of the initial voiceprint signals by using adaptive filtering technology to obtain target voiceprint signals.

[0011] According to the video and voiceprint dual-mode power equipment monitoring method provided by the application, the process of extracting features from the target video images and target voiceprint signals to obtain equipment appearance features and equipment voiceprint features comprises: Real-time detection of abnormal visual signals such as discharge light spots, smoke, and equipment appearance damage and deformation is performed by using an image recognition algorithm based on deep learning, appearance anomaly features and light intensity change features of the target video images are extracted, equipment appearance features are obtained, and an abnormal visual database is constructed; Frequency domain energy features and time domain mutation features of the target voiceprint signals are extracted, equipment voiceprint features are obtained, and an abnormal voiceprint database is constructed.

[0012] The application provides a video and voiceprint dual-mode power equipment monitoring method based on a convolutional neural network model, and the process of equipment anomaly detection combining equipment appearance features and equipment voiceprint features comprises the following steps: The process of equipment anomaly detection combining equipment appearance features and equipment voiceprint features based on a convolutional neural network model obtains visual anomaly signals and voiceprint anomaly signals. If the light intensity change in the visual anomaly signals exceeds a preset threshold range, then deep analysis is performed in combination with an abnormal visual database to determine whether there is a discharge light or appearance damage feature. If the frequency domain energy mutation in the voiceprint anomaly signals exceeds a preset threshold range, then comparison is performed in combination with an abnormal voiceprint database to determine whether there is a discharge sound or mechanical abnormal sound feature.

[0013] The application provides a video and voiceprint dual-mode power equipment monitoring method, and the process of starting dual-mode data cross verification if single-mode detection detects an anomaly comprises the following steps: Data fusion and cross verification are performed on the visual anomaly signals and the voiceprint anomaly signals to obtain multi-modal features. Multi-modal feature integration technology is used to integrate the confidence weight of the visual anomaly signals and the voiceprint anomaly signals to obtain a fused fault determination result.

[0014] The application provides a video and voiceprint dual-mode power equipment monitoring method, and the process of data fusion and cross verification on the visual anomaly signals and the voiceprint anomaly signals comprises the following steps: If the equipment appearance features match at least one abnormal visual signal in the abnormal visual database, then an abnormal visual state is determined. By comparing the light intensity change feature with a preset light intensity threshold, the light intensity anomaly degree is determined to obtain an abnormal visual level. If the equipment voiceprint features match at least one abnormal voiceprint signal in the abnormal voiceprint database, then an abnormal voiceprint state is determined. By comparing the frequency domain energy feature with a preset energy threshold, the voiceprint anomaly degree is determined to obtain an abnormal voiceprint level. According to the abnormal visual level and the abnormal voiceprint level, a weighted fusion algorithm is used to calculate a comprehensive anomaly score to obtain an equipment anomaly state. Feature data in a historical abnormal visual database and an abnormal voiceprint database are acquired, a support vector machine algorithm is used to classify the comprehensive anomaly score, and an equipment anomaly type is determined. Through time series analysis, an anomaly occurrence frequency and a change trend are extracted from the equipment anomaly type to obtain an equipment operation state prediction result.

[0015] According to the video and voiceprint dual-mode power equipment monitoring method provided by the application, the process of saving an abnormal event complete data packet and synchronously uploading to the cloud comprises: According to the fusion fault determination result, the event data determined as abnormal is locally stored, the corresponding video segment and voiceprint segment are saved, and timestamp information is attached to generate a complete event data packet; The complete event data packet is transmitted to the cloud platform through an encryption communication protocol, data is classified and archived by combining remote storage technology, and a cloud storage path is obtained for subsequent review and analysis.

[0016] According to the video and voiceprint dual-mode power equipment monitoring method provided by the application, the process of saving an abnormal event complete data packet and synchronously uploading to the cloud comprises: According to the fault severity, sound and light alarms are triggered, SMS / email notifications are sent, and a maintenance work order is generated and pushed to the operation and maintenance system.

[0017] In another aspect, the application also provides a video and voiceprint dual-mode power equipment monitoring device, which comprises: A data acquisition module is configured to synchronously acquire initial video images and initial voiceprint signals of power equipment operation, and to pre-process the initial video images and initial voiceprint signals to obtain target video images and target voiceprint signals; A feature extraction module is configured to extract features from the target video images and target voiceprint signals to obtain equipment appearance features and equipment voiceprint features; An anomaly detection module is configured to perform equipment anomaly detection based on a convolutional neural network model in combination with the equipment appearance features and equipment voiceprint features, to start dual-mode data cross-validation if an anomaly is detected by a single mode, and to trigger a high-confidence warning if an anomaly is detected by both modes; A data transmission module is configured to save an abnormal event complete data packet and synchronously upload to the cloud; A warning notification module is configured to send a warning message and notify relevant personnel through a hierarchical alarm mechanism.

[0018] In another aspect, the application also provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the video and voiceprint dual-mode power equipment monitoring method according to any of the above when executing the program.

[0019] In another aspect, the application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the video and voiceprint dual-mode power equipment monitoring method according to any of the above.

[0020] In another aspect, the present application also provides a computer program product comprising a computer program which, when executed by a processor, implements any of the above-mentioned video and voiceprint dual-modal power equipment monitoring methods.

[0021] The video and voiceprint dual-modal power equipment monitoring method and device provided by the present application synchronously collect video images and voiceprint signals of equipment operation, pre-process and extract features therefrom to obtain equipment appearance features and voiceprint features. Convolutional neural network models are used to combine the two features for anomaly detection. If a single mode detects an anomaly, dual-modal cross-validation is started. If both modes detect an anomaly, a high-confidence early warning is triggered. The present application also saves complete data packets of abnormal events and uploads them to the cloud to send a hierarchical early warning to relevant personnel. This method improves the accuracy and reliability of power equipment anomaly detection through visual and acoustic dual-modal fusion analysis, realizes real-time monitoring and timely early warning of equipment operating conditions, effectively reduces the risk of equipment failure, and improves the safe operation level of the power system.

[0022] Meanwhile, the present application can detect discharge sparks, abnormal sounds and other fault features of electrical equipment in real time, and realize precise early warning and data storage and uploading functions through multi-modal data analysis. Through dual-modal cross-validation, the false alarm rate is significantly reduced, and cloud data supports machine learning model iteration and optimization, improving long-term monitoring performance. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0024] Figure 1 is a flowchart of the video and voiceprint dual-modal power equipment monitoring method provided by the embodiment of the present application; Figure 2 is a structural schematic diagram of the video and voiceprint dual-modal power equipment monitoring device provided by the embodiment of the present application; Figure 3 is a structural schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0026] Figure 1 It is a flow chart of the video and voiceprint dual-modal power equipment monitoring method provided by an embodiment of the present invention.

[0027] like Figure 1 As shown, the execution subject of the video and voiceprint dual-modal power equipment monitoring method provided in the embodiment of the present invention can be an electronic device, and the method mainly includes the following steps: 101. Synchronously collect an initial video image and an initial voiceprint signal of the power equipment operation, and pre-process the initial video image and the initial voiceprint signal to obtain a target video image and a target voiceprint signal; In a specific implementation process, when monitoring the operating status of power equipment, the initial video image is first captured by a high-definition camera at a rate of 30 frames per second. At the same time, a high-sensitivity microphone is used to synchronously record the initial voiceprint signal at a sampling rate of 44.1kHz. The video image is denoised, and the median filter algorithm is used to remove salt and pepper noise in the image. The brightness contrast is adjusted to the standard range (the brightness mean is adjusted to 128, and the standard deviation is controlled within 20); the fast Fourier transform (FFT) is applied to the voiceprint signal to remove background noise, and the effective signal in the frequency range of 20Hz to 8kHz is extracted to obtain the target video image and target voiceprint signal.

[0028] 102. Extract features from the target video image and target voiceprint signal to obtain device appearance features and device voiceprint features; In a specific implementation process, the appearance features of the target video image can be extracted based on the deep learning framework, and the pre-trained ResNet-50 model can be used to output a 2048-dimensional feature vector, focusing on abnormalities such as cracks and deformation on the device surface. The target voiceprint signal is extracted through Mel-frequency cepstral coefficients (MFCC) to generate a 39-dimensional feature vector to analyze the vibration frequency abnormalities during device operation (such as exceeding the normal range by ±10%).

[0029] 103. Based on the convolutional neural network model, device anomaly detection is performed by combining device appearance features and device voiceprint features. If an anomaly is detected by a single modality, a dual-modal data cross-validation is initiated; if an anomaly is detected by both modalities, a high-confidence warning is triggered; 104. Save the complete data packet of the abnormal event and upload it to the cloud synchronously; In one specific implementation process, a convolutional neural network (CNN) model can be constructed, the appearance features and the voiceprint features are respectively input into two branch networks, the number of hidden layer neurons is set to 512, the ReLU activation function is adopted, the abnormal probability is output, if the single modal abnormal probability exceeds 0.8, the double modal cross verification is started, the weighted average value (the weight is 0.6 and 0.4) of the two modal abnormal probabilities is calculated, if both exceed 0.85, it is confirmed as high confidence anomaly and triggers the early warning. The abnormal event data packet will save the video in H.264 encoding format and the audio in WAV format, the data packet size is controlled within 50MB, and is synchronously uploaded to the cloud server through the HTTPS protocol, ensuring that the data transmission rate is not less than 10Mbps.

[0030] 105, issuing a warning information, and notifying relevant personnel through a hierarchical alarm mechanism.

[0031] In one specific implementation process, the system automatically generates warning information, based on abnormal severity grading (probability 0.85-0.9 for yellow warning, 0.9 or above for red warning), pushes it to the relevant device management system through the API interface, and synchronously records the log for subsequent analysis, ensuring the full process automation and real-time from data collection to warning.

[0032] The video and voiceprint double modal power equipment monitoring method of the embodiment, by synchronously collecting video images and voiceprint signals of the equipment running, pre-processing and feature extraction are performed, obtaining the equipment appearance features and voiceprint features. Convolutional neural network model is used to combine the two features for anomaly detection, if single modal detects anomaly, double modal cross verification is started, and high confidence warning is triggered when both modalities detect anomaly. The application also saves the complete data packet of the abnormal event and uploads it to the cloud, and issues a hierarchical warning to notify relevant personnel. This method combines visual and audio modalities for analysis, improves the accuracy and reliability of power equipment anomaly detection, realizes real-time monitoring and timely warning of equipment running state, effectively reduces the risk of equipment failure, and improves the safe operation level of the power system.

[0033] In one specific implementation process, the process of obtaining target video images and target voiceprint signals includes: The running data of the power equipment is synchronously collected by a high-definition camera and a microphone array, the device appearance video images and sound signals are recorded in real time, the initial video images and initial voiceprint signals are obtained; The video frames in the initial video images are segmented and enhanced by image processing technology, and the target video images are obtained; The environmental noise of the initial voiceprint signal is separated and suppressed by adaptive filtering technology, and the target voiceprint signal is obtained.

[0034] Further, the embodiment synchronously collects data of the power equipment through a high-definition camera and a microphone array, records appearance videos and sound signals in real time, and obtains initial video images and initial voiceprint signals. Image processing technology is used to segment the video frames in the initial video images, separate the key area images, and obtain segmented video frames. Enhancement processing technology is used to optimize the details of the segmented video frames, improve the image definition, and obtain target video images. According to the feature distribution of the target video images, the texture information of the key area is extracted, and the abnormal features of the equipment appearance are determined. Adaptive filtering technology is used to separate the environmental noise of the initial voiceprint signal, suppress irrelevant interference, and obtain a target voiceprint signal. According to the spectral characteristics of the target voiceprint signal, the abnormal frequency components in the sound signal are analyzed, and potential problems of the equipment running state are judged. A pre-established classification model is used to comprehensively analyze the abnormal features of the target video images and the abnormal frequency components of the target voiceprint signal, and a comprehensive evaluation result of the equipment running state is obtained.

[0035] In one specific implementation process, in the process of obtaining the target video image and the target voiceprint signal, first, the running data of the power equipment is synchronously collected through a high-definition camera and a microphone array. Specifically, a camera with a resolution of 1920x1080 is used to record the appearance video of the equipment at a rate of 30 frames per second, and an 8-channel microphone array is used to collect sound signals at a sampling rate of 44.1 kHz, ensuring that the time synchronization error of video and audio data is less than 10 milliseconds, and data matching is realized through timestamp alignment technology, laying a foundation for subsequent analysis. Then, the initial video image is processed, the image segmentation algorithm in the OpenCV library is used to divide the video frames into equipment main body and background according to the area, the Canny edge detection algorithm is used to extract the equipment contour, the threshold is set to 100 and 200 to filter noise, then the image brightness is adjusted to 1.2 times through contrast enhancement technology to improve the visibility of details, and finally the target video image is generated for defect recognition. Next, for the environmental noise separation of the initial voiceprint signal, adaptive filtering technology such as LMS algorithm is applied, the filter order is set to 32, the step parameter is set to 0.01, the noise component is iteratively calculated and subtracted from the original signal, and the analysis shows that the signal-to-noise ratio is improved to more than 15 dB after noise suppression, while the core frequency characteristics of the equipment running are preserved within the range of 200 Hz to 5 kHz, and the clear target voiceprint signal is obtained. In order to ensure the logical correlation, the processed video and voiceprint data are jointly analyzed through feature matching algorithm, for example, the vibration frequency of the equipment in the video is compared with the spectral characteristics of the voiceprint signal, and when the correlation coefficient is more than 0.85, the data consistency is confirmed, providing a reliable basis for subsequent equipment state evaluation.

[0036] In a specific implementation process, the feature extraction of the target video image and the target voiceprint signal includes: An image recognition algorithm based on deep learning is used to detect abnormal visual signals such as discharge light spots, smoke, and device appearance damage and deformation in real time, extract appearance abnormal features and light intensity change features of the target video image, obtain device appearance features, and construct an abnormal visual database. The frequency domain energy features and time domain mutation features of the target voiceprint signal are extracted, the device voiceprint features are obtained, and an abnormal voiceprint database is constructed.

[0037] For example, the feature extraction process for the target video image and the target voiceprint signal can be automatically processed by the following specific implementation method. For feature extraction of the target video image, first, a YOLOv5 algorithm based on deep learning is used for real-time detection, the image resolution is set to 1920x1080 pixels, the frame rate is set to 30 frames per second, and abnormal visual signals such as discharge light spots, smoke, and device appearance damage and deformation are detected. Through a pre-trained model, the image is detected, the confidence threshold is set to 0.85, the appearance abnormal features such as damage area ratio (for example, the detected damage area accounts for 15% of the total area of the device) and light intensity change features (for example, the light intensity value changes from 1000 lumens to 5000 lumens within 10 seconds) are extracted, thereby constructing a device appearance feature dataset, and storing the abnormal data into an abnormal visual database, which is indexed by time stamp and abnormal type (such as "light spot anomaly"), facilitating subsequent analysis. Then, for feature extraction of the target voiceprint signal, a short-time Fourier transform (STFT) algorithm is used for frequency domain analysis of the audio signal, the sampling rate is 44.1 kHz, the window length is 1024 points, the frequency domain energy features are extracted, for example, when the energy peak value in the 2kHz to 5kHz frequency band reaches 80 decibels, it is marked as abnormal; at the same time, through time domain analysis, the mutation feature is detected, the mutation threshold is set to a signal amplitude change of more than 50% within 0.1 seconds, thereby obtaining device voiceprint features, and storing the abnormal data (such as mutation time point and corresponding amplitude) into an abnormal voiceprint database, which is classified according to frequency band and mutation feature. To form a logical chain, the abnormal visual database and the abnormal voiceprint database are analyzed, for example, when the visual detects a light intensity mutation and the voiceprint signal has a 2kHz high energy peak, the system automatically marks it as a "discharge anomaly" event, and records the relevant feature values for subsequent fault diagnosis, ensuring closed-loop processing of feature extraction and anomaly recognition.

[0038] In a specific implementation process, based on a convolutional neural network model, the device anomaly detection process combining device appearance features and device voiceprint features includes: Based on the convolutional neural network model, device anomaly detection is performed by combining device appearance features and device voiceprint features to obtain visual anomaly signals and voiceprint anomaly signals; If the light intensity change in the visual abnormality signal exceeds the preset threshold range, in-depth analysis will be conducted in conjunction with the abnormal visual database to determine whether there is discharge light or appearance damage characteristics; If the frequency domain energy mutation in the abnormal voiceprint signal exceeds the preset threshold range, it will be compared with the abnormal voiceprint database to determine whether there are discharge sounds or mechanical abnormal noise characteristics.

[0039] For example, a convolutional neural network model is used to detect device anomalies. First, a camera captures an image of the device's appearance, obtaining an RGB image with a resolution of 1920×1080. This image is then fed into a pre-trained ResNet-50 convolutional neural network to extract a 2048-dimensional appearance feature vector. Simultaneously, a microphone array collects the device's voiceprint signal during operation at a sampling rate of 44.1kHz for 10 seconds. The resulting time-domain waveform is converted to a spectrogram using a short-time Fourier transform (STFT) with a resolution of 256×256. This is then fed into a customized VGG-16 network to extract a 512-dimensional voiceprint feature vector. After fusing the two features, a fully connected layer is used to reduce the feature vector's dimensionality to 128. This is then fed into a Softmax classifier, which outputs anomaly probabilities for both the visual and voiceprint anomaly signals, with a threshold of 0.9. During visual anomaly signal detection, if the light intensity variation exceeds a preset threshold, the feature vector is detected.

[0040] For example, if the average brightness changes by more than 50 cd / m², the image is compared with an abnormal visual database (including samples of discharge flash and surface damage) and the cosine similarity algorithm is used to calculate the similarity of the feature vectors. If the similarity is greater than 0.95, it is determined to be a discharge flash or surface damage. In voiceprint abnormal signal detection, if the frequency domain energy mutation exceeds the threshold.

[0041] For example, if the energy in a specific frequency band exceeds 80dB, the spectrum is compared with an abnormal voiceprint database (containing samples of discharge sounds and mechanical noises). The dynamic time warping (DTW) algorithm is used to calculate the spectral sequence distance. If the distance is less than a preset threshold of 10, it is determined to be an abnormal voiceprint. Finally, the two types of abnormal signals are combined and weighted using a logistic regression model (visual weight 0.6, voiceprint weight 0.4) to output a comprehensive abnormality probability. If it exceeds 0.85, a device abnormality alarm is triggered. This process forms a closed-loop logic through feature extraction, threshold judgment, and database comparison to ensure detection accuracy.

[0042] In a specific implementation, if a single modality detects an anomaly, the process of initiating bimodal data cross-validation includes: Data fusion and cross-validation are performed on the visual anomaly signal and the voiceprint anomaly signal to obtain a multi-modal feature; The confidence weight of the visual anomaly signal and the voiceprint anomaly signal is integrated by using a multi-modal feature integration technology to obtain a fused fault determination result.

[0043] For example, when a single modality detects an anomaly, for example, the visual modality detects an abnormal device surface temperature through a convolutional neural network, triggering a dual-modality data cross-validation. First, the visual modality extracts features from the infrared thermal image, processes the image using a VGG-16 model to obtain a 256-dimensional feature vector, and detects that the temperature value of the abnormal area reaches 85.3°C, exceeding the normal threshold of 75°C. At the same time, the voiceprint modality collects device operation sound through a microphone array, extracts features through a mel-frequency cepstral coefficient, generates a 128-dimensional feature vector, and detects that the frequency abnormal peak value is 2.5kHz, exceeding the normal range of 1.8kHz. In the data fusion stage, principal component analysis is used to reduce the visual and voiceprint features to 64 dimensions, retaining 95% of the variance to ensure minimal information loss. Subsequently, Kalman filtering is used to smooth the time series data of the two modalities, eliminating noise interference, and obtaining a fused multi-modal feature vector. Then, a multi-modal feature integration technology is used to integrate the confidence by using a weighted average method, with the visual modality confidence being 0.85 (based on the significance of the temperature anomaly) and the voiceprint modality confidence being 0.65 (due to the small amplitude of the frequency anomaly). According to the weight calculation formula, the fused confidence is 0.763. If the confidence threshold is set to 0.7, it is determined to be a fault. Finally, the system automatically records the fault type as "high temperature accompanied by vibration anomaly" and pushes the result to a maintenance scheduling system, triggering a repair task, logically ensuring the full-process automation from anomaly detection to fault determination.

[0044] In one specific implementation process, the process of data fusion and cross-validation of the visual anomaly signal and the voiceprint anomaly signal includes: If the device appearance feature matches at least one abnormal visual signal in the abnormal visual database, the abnormal visual state is determined; By comparing the light intensity change feature with the preset light intensity threshold, the light intensity anomaly degree is determined to obtain an abnormal visual level; If the device voiceprint feature matches at least one abnormal voiceprint signal in the abnormal voiceprint database, the abnormal voiceprint state is determined; By comparing the frequency energy feature with the preset energy threshold, the voiceprint anomaly degree is determined to obtain an abnormal voiceprint level; According to the abnormal visual level and the abnormal voiceprint level, a weighted fusion algorithm is used to calculate a comprehensive anomaly score to obtain a device anomaly state; Obtain feature data in the historical abnormal visual database and the abnormal voiceprint database, classify the comprehensive abnormal score by using a support vector machine algorithm, and determine the device abnormal type; Extract the abnormal occurrence frequency and the change trend from the device abnormal type by time series analysis, and obtain the device running state prediction result.

[0045] In a specific implementation process, the process of saving the complete data packet of the abnormal event and synchronously uploading to the cloud includes: According to the fused fault determination result, the event data determined as abnormal is stored locally, the corresponding video segment and voiceprint segment are saved, and the timestamp information is attached, and a complete event data packet is generated; The complete event data packet is transmitted to the cloud platform through an encryption communication protocol, the data is classified and archived by using a remote storage technology, and a cloud storage path is obtained for subsequent retrieval and analysis.

[0046] Specifically, for the fault determination result, the event data determined as abnormal is obtained, the corresponding video segment and voiceprint segment are extracted therefrom, the timestamp information is attached, and a complete event data packet is generated. According to the generated event data packet, the local storage technology is used to save the event data packet to a preset storage area, the metadata information of the data packet is synchronously recorded, and identification information of storage completion is obtained. Through the identification information of storage completion, an encryption communication protocol is triggered to encrypt the event data packet, and an encrypted event data packet is generated. For the encrypted event data packet, a secure transmission channel is used to upload the event data packet to the cloud platform, and confirmation information of successful uploading is obtained. According to the confirmation information of successful uploading, the encrypted event data packet is classified and archived on the cloud platform by using a remote storage technology, and an archived storage path is determined. Through the archived storage path, an index record available for retrieval is generated and saved to the cloud database, and a queryable path identifier is obtained. If a query request is triggered, the corresponding storage path is extracted from the cloud database according to the path identifier, the related event data packet is obtained, and the data retrieval process is completed.

[0047] In a specific implementation process, the process of issuing a warning information and notifying relevant personnel through a hierarchical alarm mechanism includes: According to the fault severity, an audible and light alarm, an SMS / email notification, and a maintenance work order pushed to an operation and maintenance system are triggered.

[0048] For example, in the process of issuing early warning information and notifying relevant personnel through a hierarchical alarm mechanism, first, an audible and light alarm is triggered according to the severity of the fault. The impact range and device importance can be quantitatively analyzed by the built-in fault evaluation algorithm of the system. For example, a fault that affects more than 1000 users or a core device downtime of more than 5 minutes is defined as a first-level fault. When the impact coefficient is 0.8 or higher, a red audible and light alarm is triggered, the alarm volume is set to 85 decibels, the light flashing frequency is 2 times per second, and the alarm triggering time and fault details are recorded to the log database to form a traceable data chain for subsequent analysis of fault frequency and response efficiency. Next, the system automatically sends a short message or email notification based on the hierarchical results. For example, in the case of a first-level fault, the system calls the SMS gateway interface and sends an emergency notification containing the fault location, impact range, and estimated recovery time (such as within 2 hours) to the pre-set 10 core operation and maintenance personnel within 30 seconds, and sends an email to the second-level management personnel. The email content includes a preliminary analysis report of the fault, which includes fault occurrence rate statistics, such as 3 times of the same fault in the past 24 hours, and trigger cause proportion analysis for hardware problems accounting for 60%, to facilitate quick decision-making by management. Finally, the system automatically generates a maintenance work order and pushes it to the operation and maintenance system. For example, through the API interface, the work order information (such as fault number, priority level 1, and required response time within 30 minutes) is synchronized to the operation and maintenance platform, and the fault diagnosis data such as abnormal device operating temperature rise to 75 degrees Celsius, exceeding the normal range by 15%, is attached to the work order. Combined with historical maintenance records, it is analyzed that the possible cause is a cooling system failure. After pushing, the system monitors the work order status in real time, and if it is not responded within 2 hours, it is automatically upgraded to a higher management level to ensure the fault handling closed loop. Through the above process, the early warning information is processed in a hierarchical manner, and each link is data-driven, forming a full-link automated management from alarm to work order.

[0049] Based on the same general inventive concept, the present application also protects a video and voiceprint dual-mode power equipment monitoring device. The video and voiceprint dual-mode power equipment monitoring device provided by the present application is described below, and the video and voiceprint dual-mode power equipment monitoring device described below can be mutually corresponding to the video and voiceprint dual-mode power equipment monitoring method described above.

[0050] Figure 2 is a structural schematic diagram of the video and voiceprint dual-mode power equipment monitoring device provided by the embodiment of the present application, as Figure 2 indicated, the video and voiceprint dual-mode power equipment monitoring device of the embodiment includes a data acquisition module 21, a feature extraction module 22, an anomaly detection module 23, a data transmission module 24, and a warning notification module 25.

[0051] The data acquisition module 21 is configured to synchronously acquire an initial video image and an initial voiceprint signal of the operation of the power equipment, and pre-process the initial video image and the initial voiceprint signal to obtain a target video image and a target voiceprint signal. The feature extraction module 22 is configured to extract features from the target video image and the target voiceprint signal to obtain equipment appearance features and equipment voiceprint features. The anomaly detection module 23 is configured to perform equipment anomaly detection based on a convolutional neural network model in combination with the equipment appearance features and the equipment voiceprint features, if an anomaly is detected by a single mode, cross-verification of dual-mode data is started, and if anomalies are detected by both dual modes, a high-confidence early warning is triggered. The data transmission module 24 is configured to save a complete data packet of an abnormal event and synchronously upload the complete data packet to a cloud. The early warning notification module 25 is configured to issue early warning information and notify relevant personnel through a hierarchical alarm mechanism.

[0052] Figure 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. The video and voiceprint dual-mode power equipment monitoring device can include a processor 310, a communications interface 320, a memory 330 and a communications bus 340, wherein the processor 310, the communications interface 320 and the memory 330 complete mutual communication through the communications bus 340. The processor 310 can invoke logical instructions in the memory 330 to execute a video and voiceprint dual-mode power equipment monitoring method.

[0053] In addition, the logical instructions in the memory 330 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0054] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program being stored on a non-transitory computer readable storage medium, and the computer program being capable of executing the video and voiceprint dual-modal power equipment monitoring method provided by the above-mentioned methods when executed by a processor.

[0055] In yet another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is capable of implementing the video and voiceprint dual-modal power equipment monitoring method provided by the above-mentioned methods when executed by a processor.

[0056] It should be noted that the related information involved in each embodiment of the present application is strictly in accordance with the requirements of laws and regulations, and follows the principles of legality, legitimacy and necessity, and is based on the legitimate purpose of the business scene, and processes the information provided by the user in the process of using the product / service or generated due to the use of the product / service, and the information authorized by the user.

[0057] The related information processed by the present application will be different due to the specific product / service scene, and the specific scene of the user using the product / service should be used as the standard, which may involve the user's account information, device information or other related information. The present application will treat the related information and its processing with high diligence.

[0058] The present application attaches great importance to the security of the related information, and has taken security protection measures in accordance with the industry standards, which are reasonable and feasible to protect the related information, so as to prevent the related information from being accessed, disclosed, used, modified, damaged or lost without authorization.

[0059] The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0060] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0061] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A video and voiceprint dual-modal power equipment monitoring method, characterized in that: include: Synchronously collecting an initial video image and an initial voiceprint signal of the power equipment operation, and preprocessing the initial video image and the initial voiceprint signal respectively to obtain a target video image and a target voiceprint signal; Performing feature extraction on the target video image and target voiceprint signal to obtain device appearance features and device voiceprint features; Based on the convolutional neural network model, device anomaly detection is performed by combining device appearance features and device voiceprint features. If a single modality detects an anomaly, bimodal data cross-validation is initiated; if both modalities detect an anomaly, a high-confidence warning is triggered. Save the complete data package of abnormal events and upload it to the cloud synchronously; Issue early warning information and notify relevant personnel through a hierarchical alarm mechanism.

2. The video and voiceprint dual-modal power equipment monitoring method according to claim 1 is characterized in that: The process of obtaining the target video image and target voiceprint signal includes: The operating data of power equipment is collected synchronously through high-definition cameras and microphone arrays, and the video images and sound signals of the equipment appearance are recorded in real time to obtain the initial video images and initial voiceprint signals; Using image processing technology to segment and enhance the video frames in the initial video image to obtain a target video image; The environmental noise of the initial voiceprint signal is separated and suppressed by adaptive filtering technology to obtain a target voiceprint signal.

3. The video and voiceprint dual-modal power equipment monitoring method according to claim 1 is characterized in that: The process of extracting features from the target video image and target voiceprint signal to obtain device appearance features and device voiceprint features includes: Use a deep learning-based image recognition algorithm to detect abnormal visual signals such as discharge spots, smoke, and equipment damage and deformation in real time, extract the abnormal appearance features and light intensity change features of the target video image, obtain the equipment appearance features and build an abnormal visual database; The frequency domain energy features and time domain mutation features of the target voiceprint signal are extracted to obtain the device voiceprint features and construct an abnormal voiceprint database.

4. The video and voiceprint dual-modal power equipment monitoring method according to claim 1 is characterized in that: The process of detecting device anomalies based on a convolutional neural network model and combining device appearance features and device voiceprint features includes: Based on the convolutional neural network model, device anomaly detection is performed by combining device appearance features and device voiceprint features to obtain visual anomaly signals and voiceprint anomaly signals; If the light intensity change in the visual abnormality signal exceeds a preset threshold range, an in-depth analysis is performed in conjunction with the abnormal visual database to determine whether there is a discharge flash or appearance damage feature; If the frequency domain energy mutation in the abnormal voiceprint signal exceeds a preset threshold range, it is compared with the abnormal voiceprint database to determine whether there are discharge sound or mechanical abnormal sound characteristics.

5. The video and voiceprint dual-modal power equipment monitoring method according to claim 1 is characterized in that: If an anomaly is detected in a single modality, the process of initiating bimodal data cross-validation includes: Perform data fusion and cross-validation on visual abnormal signals and voiceprint abnormal signals to obtain multimodal features; Multimodal feature integration technology is used to integrate the confidence weights of visual abnormal signals and voiceprint abnormal signals to obtain the fused fault judgment results.

6. The video and voiceprint dual-modal power equipment monitoring method according to claim 5 is characterized in that: The process of data fusion and cross-validation of visual anomaly signals and voiceprint anomaly signals includes: If the device appearance feature matches at least one abnormal visual signal in the abnormal visual database, an abnormal visual state is determined; By comparing the light intensity change characteristics with the preset light intensity threshold, the abnormal degree of light intensity is judged and the abnormal visual level is obtained; If the device voiceprint feature matches at least one abnormal voiceprint signal in the abnormal voiceprint database, the abnormal voiceprint status is determined; By comparing the frequency domain energy characteristics with the preset energy threshold, the abnormality of the voiceprint is determined and the abnormal voiceprint level is obtained; According to the abnormal visual level and abnormal voiceprint level, a weighted fusion algorithm is used to calculate a comprehensive abnormality score to obtain the abnormal status of the device; Acquire feature data from a historical abnormal visual database and an abnormal voiceprint database, use a support vector machine algorithm to classify the comprehensive abnormality score, and determine the type of device abnormality; Through time series analysis, the frequency of abnormal occurrence and the changing trend are extracted from the equipment abnormality types to obtain the equipment operation status prediction result.

7. The video and voiceprint dual-modal power equipment monitoring method according to claim 1 is characterized in that: The process of saving the complete data package of an abnormal event and synchronously uploading it to the cloud includes: Based on the fused fault determination results, the event data determined to be abnormal is stored locally, the corresponding video clips and voiceprint clips are saved, and timestamp information is added to generate a complete event data packet; The complete event data packet is transmitted to the cloud platform through an encrypted communication protocol, the data is classified and archived in combination with remote storage technology, and the cloud storage path is obtained for subsequent review and analysis.

8. The video and voiceprint dual-modal power equipment monitoring method according to claim 1 is characterized in that: The process of issuing early warning information and notifying relevant personnel through the hierarchical alarm mechanism includes: According to the severity of the fault, audio and visual alarms, SMS / email notifications are triggered, and a maintenance work order is generated and pushed to the operation and maintenance system.

9. A video and voiceprint dual-modal power equipment monitoring device, characterized in that: include: A data acquisition module is used to synchronously collect an initial video image and an initial voiceprint signal of the power equipment operation, and pre-process the initial video image and the initial voiceprint signal to obtain a target video image and a target voiceprint signal; A feature extraction module is used to extract features from the target video image and target voiceprint signal to obtain device appearance features and device voiceprint features; The anomaly detection module is used to detect device anomalies based on a convolutional neural network model, combining device appearance features and device voiceprint features. If a single modality detects an anomaly, a dual-modal data cross-validation is initiated; if both modalities detect an anomaly, a high-confidence warning is triggered. Data transmission module, used to save the complete data packet of abnormal events and upload it to the cloud synchronously; The early warning notification module is used to issue early warning information and notify relevant personnel through a hierarchical alarm mechanism.

10. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the video and voiceprint dual-modal power equipment monitoring method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Comprehensive monitoring and management system for transformer substation

    CN121440928A

  • Power equipment fault marking method based on waveform and unmanned aerial vehicle video fusion

    CN122289993A