Power violation operation identification method based on multi-modal fusion
The multi-modal data fusion method using neural networks addresses inefficiencies and inaccuracies in traditional power violation detection by synchronizing and analyzing data from various perspectives, improving detection accuracy and efficiency.
Patent Information
- Application Number
- CN202510366047.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-15
AI Technical Summary
Traditional power violation identification methods are inefficient and have low accuracy, which cannot meet the needs of industry development. Manual inspection and basic monitoring equipment lacks automated analysis functions, making it difficult to quickly adapt to changes in power operation scenarios and identify complex violations.
Through the multimodal fusion method, image data and environmental equipment data of the power operation site are collected from multiple perspectives and different distances, deep convolutional neural networks are used for feature extraction and fusion, and dynamic recognition is performed with neural networks with timing encoding capabilities, so as to achieve rapid judgment of violation operations.
It improves the accuracy and efficiency of power violation identification, reduces the rate of misjudgment, enhances the safety and risk management level of the power operation site, and can accurately identify violations under complex lighting and occlusion.
Smart Images

Figure CN120318647A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system monitoring. Specifically, it relates to a method for identifying power violation operations based on multi-modal fusion. Background Art
[0002] In today's society, electricity, as a core energy source, the stable operation of its system is crucial for the development of the national economy and people's livelihood. Moreover, with the acceleration of the urbanization and industrialization processes, the complexity and frequency of power operations have increased significantly. Currently, power violation identification mainly relies on manual inspections and basic monitoring, and this method faces many problems. Among them, manual inspections are inefficient, it is difficult to achieve high-frequency full coverage, inspection personnel are prone to fatigue and miss risk points, while basic monitoring devices lack automated analysis functions, and a large number of videos need to be screened manually one by one, making it difficult to capture violation behaviors in a timely manner. Also, manual inspections are limited by the professional qualities and vision blind spots of personnel, making it difficult to accurately identify complex violation operations. The image quality of traditional monitoring systems is also poor under complex lighting and occlusion conditions, resulting in low recognition rates. At the same time, power operation scenarios often change, and existing monitoring methods are difficult to quickly adapt to new technologies and risk points, lacking flexibility and often requiring long-time adjustments. Currently, most power operation data is simply stored and not deeply utilized, lacking intelligent decision-making support, and unable to quickly and effectively judge the risk level and provide data support for preventive measures.
[0003] Existing Technical Problems: Traditional power violation identification means are not only inefficient but also inaccurate, unable to meet the needs of industry development. Summary of the Invention
[0004] The present invention solves the technical problem that traditional power violation identification means are not only inefficient but also inaccurate, unable to meet the needs of industry development. The present invention collects image data of the power operation site from multiple perspectives and different distances, as well as obtains environmental and equipment data, and realizes the rapid determination of violation operations through neural network algorithms, effectively improving the recognition accuracy and recognition efficiency.
[0005] To solve the above problems, the present invention provides a method for identifying electric power violation operations based on multi-modal fusion, including: collecting electric power operation data, collecting image data of the electric power operation site from multiple perspectives and different distances, and obtaining environmental and equipment data; constructing multi-modal data, pairing the data of different modalities captured at the same moment in the data of the electric power operation site obtained according to the time stamp; feature adaptation, using a deep convolutional neural network to align the visual semantics of objects, scenes, and actions in the image data with the semantics of words and sentences in the language, and combining the environmental and equipment data to form feature data; feature pruning and injection, performing multi-layer deep fusion and pruning optimization on the feature data to screen out key features; dynamic identification and determination, based on a neural network with time series encoding ability, dynamically identifying the key features input to the neural network with time series encoding ability in the order of time stamp, and outputting a determination result on whether the electric power operation is a violation operation.
[0006] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By collecting image data from multiple perspectives and different distances and combining environmental and equipment data, the effective fusion of multi-modal data is realized, which provides more comprehensive information and improves the accuracy and robustness of identification. Moreover, using the time stamp to pair different modality data enables the data captured at the same moment to be analyzed synchronously. This time series analysis ability can better capture the dynamic characteristics of electric power operations and improve the timeliness of the model's understanding and judgment of operation behaviors. At the same time, using a deep convolutional neural network to align the image data and language semantics enables visual information and semantic information to complement each other and form a richer feature expression. This process improves the model's ability to identify complex operation behaviors. By performing multi-layer deep fusion and pruning optimization on the feature data, key features can be effectively screened out, redundant information can be reduced, and the processing efficiency and response speed of the model can be improved. This helps to achieve a faster determination speed in practical applications. Introducing a neural network with time series encoding ability enables the model to process time series features and dynamically identify the violation operations of electric power operations. This mechanism effectively improves the model's ability to control and judge the electric power operation process. Through comprehensive data collection and in-depth analysis, the misjudgment rate of identifying electric power violation operations can be significantly reduced, the safety and risk management level of the electric power operation site can be improved, and it helps to build a more reliable electric power system management mechanism.
[0007] In a possible design, collecting image data of the electric power operation site from multiple perspectives and different distances and obtaining environmental and equipment data includes: collecting visible light images of the electric power operation site from multiple perspectives and different distances, collecting heat distribution information of the electric power operation site through an infrared camera and generating infrared images; recording environmental parameters such as the illumination intensity, illumination direction, environmental temperature, and humidity of the electric power operation site; and collecting real-time operation parameters of on-site electrical equipment.
[0008] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By collecting visible light images and infrared images from multiple perspectives and at different distances, the omnidirectional monitoring of the power operation site is realized. This multi-perspective capture can reduce blind spots, ensure effective coverage of every corner of the site, and improve the integrity of scene information. Moreover, by combining visible light images and infrared images, more comprehensive information about the operation site can be provided. Among them, the heat distribution information captured by infrared images can reveal the operating status of electrical equipment, helping to identify potential problems such as overheating, which is an effective supplement to optical information. By recording environmental parameters such as light intensity, direction, environmental temperature, and humidity, the consideration of the impact of the external environment on power operations is increased. These environmental parameters can provide important background information for subsequent data analysis, helping to identify abnormal conditions and generate early warnings. Collecting the real-time operating parameters of on-site electrical equipment can help to timely understand the equipment status and quickly respond to the handling of emergencies. The real-time nature of this data provides guarantee for risk assessment and decision-making, and can enhance the safety protection ability. At the same time, by integrating multiple data sources, a richer feature set is provided for data analysis and machine learning models, which can better train the recognition system and improve the accuracy and reliability.
[0009] In a possible design, the data obtained from the power operation site is paired according to the time stamp for different modality data captured at the same moment, including: Summarizing the collected visible light images and infrared images, and pairing the different modality images captured at the same moment according to the time stamp to form multiple groups of image pairs; Performing standardization processing on the paired data, unifying the image size and resolution, and performing feature annotation.
[0010] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By using timestamps to pair different modality images captured at the same moment, the temporal consistency between different data sources is ensured. This precise temporal matching can effectively capture the comprehensive state of the power operation site at the same moment and reduce information loss caused by data imported at different time points. By pairing visible light images and infrared images to form multiple groups of image pairs, the expression ability of the power operation site state is enhanced. The infrared images provide information about temperature and heat, while the visible light images show the appearance and operation context of the equipment. This combination helps to comprehensively analyze the on-site situation. The paired data is standardized to unify the image size and resolution, ensuring the consistency of the data used in subsequent analysis. This standardization helps to reduce errors in data processing and improve the efficiency and accuracy of subsequent model training and analysis. Feature annotation is performed on the paired data to form a high-quality training set or validation set. This detailed feature annotation helps machine learning algorithms better understand the features of different modalities and their importance in power operations, thereby improving the learning effect and prediction ability of the model. At the same time, through standardization and feature annotation, the paired data is more easily utilized by subsequent analysis tools and machine learning models. The relevant algorithms for data analysis can efficiently process these annotated data to support more complex pattern recognition and prediction tasks.
[0011] In a possible design, a deep convolutional neural network is used to align the visual semantics of objects, scenes, and actions in the image data with the semantics of words and sentences in the language, and combined with environmental and device data to form feature data, including: multi-scale feature extraction, using a deep convolutional neural network to extract feature data of different scales in visible light images and infrared images; feature distribution calibration, using an adaptation module to adjust the feature data of different modalities and different scales by means of normalization and linear transformation methods to make them tend to be consistent in terms of data range and distribution form; semantic space alignment, using feature annotation information and a pre-trained semantic model to map the feature data in visible light images and infrared images to the same semantic space, so as to align the visual semantics of objects, scenes, and actions in the image data with the semantics of words and sentences in the language; feedback optimization, inputting the feature data into a small discriminant model, and reversely fine-tuning the adaptation parameters according to the discriminant results and iterating repeatedly to improve the fusion of different modality features.
[0012] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By using a deep convolutional neural network to extract feature data of different scales in visible light images and infrared images, it is possible to capture more comprehensive detailed information and overall structure in the images, and multi-scale feature extraction helps to improve the model's recognition and understanding capabilities for complex scenes and objects. The adaptation module is used to adjust the feature data of different modalities and different scales through normalization and linear transformation, making them tend to be consistent in terms of data range and distribution form. This calibration process can reduce the differences between features of different modalities, improve the effectiveness of feature fusion, and enhance the stability of the model. Moreover, by using feature annotation information and a pre-trained semantic model, the feature data in visible light images and infrared images are mapped to the same semantic space, aligning the semantics of visual information with the semantics of language information. This alignment of the semantic space enhances the model's in-depth understanding of the operation scenario and operation behavior, and improves the accurate judgment ability for power operations. At the same time, by inputting the feature data into a small discriminant model and fine-tuning the adaptation parameters in reverse according to the discriminant results, a feedback optimization mechanism is formed, enabling the model to continuously iterate and improve the feature fusion effect. This dynamic optimization mechanism can adapt to different situations and gradually improve the recognition accuracy and robustness of the model. Additionally, through the optimized feature fusion and semantic alignment, the recognition accuracy and processing efficiency of on-site operations in power work can be significantly improved. Such technical effects have important practical significance and can effectively reduce potential safety hazards and operation risks in the power field.
[0013] In a possible design, a deep convolutional neural network is used to extract feature data of different scales in visible light images and infrared images, including: at the bottom layer, obtaining fine texture features of device edges and human silhouettes; at the middle layer, capturing geometric shape features related to action postures; at the high layer, focusing on abstract concept features of action semantics and device operating states.
[0014] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By extracting underlying, middle-layer, and high-layer features layer by layer, a richer feature representation can be formed, covering multi-dimensional information from details to the whole, which can comprehensively understand the complex scenarios and their dynamic changes in the power operation site. Among them, the underlying features can accurately capture the subtle changes in the edges of equipment and the contours of personnel, ensuring high-precision recognition of objects and scenes, and providing a solid foundation for subsequent behavior recognition and anomaly detection. The middle-layer features focus on capturing the geometric shape features related to action postures, which can better understand and recognize the specific operation actions of the staff, and provide strong support for safety monitoring and operation evaluation. The high-layer features focus on the abstract concept features of action semantics and equipment operation states, enabling the system to achieve more complex reasoning, such as identifying equipment failures and evaluating operation efficiency. This grasp of abstract semantics greatly improves the effectiveness of the model in high-level decision-making and judgment. Moreover, the extraction of features at different levels enables the features of visible light images and infrared images to be effectively aligned and fused at different levels, promoting information collaboration between different modalities, thereby improving the overall performance.
[0015] In a possible design, the feature data is subjected to multi-layer deep fusion and pruning optimization to screen out key features, including: dividing the feature data into underlying, middle-layer, and high-layer levels hierarchically; injecting operations from the underlying layer to the middle layer, and using a feature importance evaluation mechanism to analyze the contribution of the underlying features to violation recognition, screening out key underlying features, and integrating them into the middle layer in a weighted fusion manner. The features after middle-layer fusion are then injected into the high layer; an adaptive fusion rule is constructed based on the correlation between the middle-layer and high-layer features, and the fusion coefficient is dynamically adjusted to promote the accurate matching of the concrete action features in the middle layer and the abstract semantics in the high layer.
[0016] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By hierarchically dividing the feature data into the underlying layer, the middle layer, and the high layer, and using the feature importance evaluation mechanism, it is possible to more accurately analyze the contribution of the underlying features to violation recognition, which can effectively remove redundant and irrelevant features and improve the data utilization efficiency. Moreover, by adopting the weighted fusion method and injecting the key underlying features into the middle-layer features, the feature fusion becomes more targeted and effective. At the same time, by fusing the fine features of the underlying layer with the action features of the middle layer, a more expressive feature combination can be formed, improving the overall recognition performance. An adaptive fusion rule is constructed based on the correlation between the middle-layer and high-layer features, and the fusion coefficient is dynamically adjusted, enabling real-time optimization of feature fusion according to the actual analysis situation, which ensures the effective combination of features of different modalities and levels and improves the model's ability to handle complex scenarios. By promoting the precise matching of the concrete action features of the middle layer with the abstract semantics of the high layer, the specific behaviors and their underlying intentions in power operations can be understood at a deeper level. This enhances the analysis and recognition of operation behaviors and makes the detection of violation behaviors more accurate. Through feature pruning optimization, the system can reduce the computational complexity to reduce the running time and resource consumption of the model, which not only makes the model more efficient but also improves the real-time monitoring and response capabilities.
[0017] In a possible design, based on a neural network with the ability of time series encoding, key features input to the neural network with the ability of time series encoding in the order of timestamps are dynamically recognized, and a determination result of whether there is a violation operation in the power operation is output, including: using a neural network with the ability of time series encoding as the basic framework to process the time series features of the power operation behavior and retain key information; inputting the key features into the neural network in the order of timestamps to enable the neural network to learn the evolution law of the action states at different time points; connecting a classifier and outputting a determination result of whether there is a violation operation in the power operation based on the integrated features.
[0018] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By using a neural network with the ability of time series encoding, the time series features of the power operation behavior can be effectively processed, which enables the model to capture the dynamically changing features, understand the operation states at different moments and their evolution, and improve the effective utilization of time series information. Moreover, inputting the key features into the network in the order of timestamps allows the neural network to learn the evolution law of the action states, which enables the model to accurately grasp the time dynamic changes in the operation process, enhances the neural network's understanding of the power operation process, and makes the model more adaptable. At the same time, by integrating the time series features and then inputting them into the classifier and deriving whether there is a violation behavior in the power operation based on the information in the time series, the accuracy of the determination result can be significantly improved, reducing the possibility of misjudgment and missed judgment.
[0019] In a possible design, the temporal characteristics of power operation behaviors are processed, and key information is retained, including: using a fully connected layer to reduce the dimension of the high-dimensional temporal feature vector, integrating and refining the key information, and reorganizing the scattered features according to the violation discrimination logic.
[0020] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By performing dimensionality reduction through the fully connected layer, the dimension of the high-dimensional temporal feature vector can be effectively reduced, which improves the processing speed of data and reduces the consumption of computing resources, thereby enhancing the operating efficiency of the model. At the same time, during the dimensionality reduction process, the model can automatically identify and extract key information, integrate scattered information, and reduce the interference of noise and redundant information. Moreover, by integrating and refining the key information, the representativeness and effectiveness of the training data are enhanced, enabling the model to better learn the patterns related to violation operations during the training process, thus improving the learning effect and the final prediction accuracy of the model. The reorganized features help the model better adapt to different power operation scenarios. Regardless of environmental changes or the diversity of operator behaviors, the model can maintain a high recognition performance.
[0021] In a possible design, the neural network with temporal encoding ability includes a long short-term memory network.
[0022] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: The long short-term memory network (LSTM) is particularly good at processing temporal data and can effectively capture the long-term dependencies in the time series. This ability gives the model a more significant advantage when analyzing power operation behaviors, enabling it to recognize and understand the impact of time delays, and thus more accurately determine whether there are violation operations. Moreover, the long short-term memory network includes a forget gate, an input gate, and an output gate, which can dynamically select to retain or forget information. This mechanism enables the network to automatically adjust the attention degree to the features at different time steps, providing higher flexibility for feature selection and avoiding the problem of gradient disappearance in traditional neural networks during long-term dependency learning.
[0023] In a possible design, the acquisition frequency of power operation data collection is adjusted according to the complexity of power operations. The higher the complexity of the power operation, the higher the acquisition frame rate, to ensure continuous capture of actions.
[0024] Compared with the prior art, the technical effects achieved by adopting this technical solution are as follows: By dynamically adjusting the acquisition frequency according to the complexity of power operations, it is possible to more accurately capture every detail of key operations. Especially in complex power operations, such as equipment maintenance, high-voltage operations, or fault troubleshooting, it can ensure that every important action is coherently recorded. Moreover, a higher acquisition frequency helps to improve the timeliness and integrity of data. Compared with the acquisition scheme with a fixed frequency, the adjusted frequency can provide higher-resolution data during high-complexity operations, reduce information loss, and improve the effectiveness of subsequent analysis and model training. In the case of low complexity of power operations, the acquisition frequency can be appropriately reduced, thereby reducing the amount of data generated and optimizing storage and computing resources. This flexible frequency adjustment mechanism can efficiently utilize resources, reduce costs, and avoid invalid data redundancy. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flowchart of a method for identifying power violation operations based on multimodal fusion provided by an embodiment of the present invention; Figure 2 It is an algorithm flowchart of a method for identifying power violation operations based on multimodal fusion provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of constructing multimodal data provided by an embodiment of the present invention; Figure 4 It is a schematic structural diagram of feature pruning and injection provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention is provided in conjunction with the accompanying drawings.
[0027] Refer to Figures 1 to 4 , this embodiment provides a method for identifying power violation operations based on multimodal fusion, including: power operation data acquisition, acquiring image data of the power operation site from multiple perspectives and different distances, and obtaining environmental and equipment data; constructing multimodal data, pairing the data obtained from the power operation site according to the time stamp for different modalities captured at the same moment; feature adaptation, using a deep convolutional neural network to align the visual semantics of objects, scenes, and actions in the image data with the semantics of words and sentences in the language, and combining environmental and equipment data to form feature data; feature pruning and injection, performing multi-layer deep fusion and pruning optimization on the feature data to screen out key features; dynamic recognition and determination, based on a neural network with the ability of temporal encoding, dynamically recognizing the key features input to the neural network with the ability of temporal encoding in the order of time stamp, and outputting a determination result of whether the power operation is a violation operation.
[0028] Specifically, the steps for collecting power operation data include: Selecting suitable image and infrared thermal imaging devices to collect visible light, infrared images, as well as environmental and equipment operation data. Constructing a synchronous acquisition mechanism, relying on a high-precision clock synchronization module, etc., to make different modality data correspond to the same moment and ensure spatio-temporal consistency. Using sensors to collect environmental data for auxiliary deviation correction, arranging special personnel to label operation tags, and providing accurate supervision information for model training. Steps for constructing multi-modal data: Summarize the visible light images and infrared images collected in the early stage, use accurate timestamps as indexes, pair images of different modalities at the same moment, and construct a closely related image group. Subsequently, carry out standardization processing to unify the sizes and resolutions of each image, and eliminate the problem of inconsistent data specifications caused by equipment differences. Use histogram equalization to enhance the images, improve the clarity and contrast of the visible light images, and enhance the display accuracy of the temperature signals of the infrared images. Based on electrical engineering expertise, further expand the feature annotations of key visual elements, sort out the data context, and build a high-quality multi-modal data architecture. Feature adaptation steps: Use a deep convolutional neural network to extract multi-scale features from multi-modal data, from shallow texture contours to deep semantic concepts. Use batch normalization technology to calibrate the feature distribution, reduce the fluctuations caused by lighting and equipment differences, and then combine with a pre-trained model to map different modality features to the same semantic space, and optimize through the feedback of the discriminant model to achieve accurate semantic alignment. Feature pruning and injection steps: First, extract features layer by layer, sorting from low-level texture to high-level semantic information. During bottom-layer fusion, according to the characteristics of the power scene, weights are assigned to the fine actions of visible light and the temperature clues of infrared, and after weighted average fusion, they are injected into the middle layer; the middle layer integrates local action structures. As the power operation process progresses, the importance of various actions at different stages changes dynamically. At the initial equipment startup stage, the action of pressing the start button by hand is crucial, and the relevant feature weights are increased; at the subsequent stable operation monitoring stage, the weights of limb actions are appropriately reduced, and the weights of equipment operation state features increase. Again, fuse the bottom-layer optimized features and the newly extracted middle-layer features, strengthen the tracking of limb dynamic changes and integrate them into the bottom-layer optimized features; the high layer contains abstract semantics, integrate the key action postures in the middle layer, and lock the nature of violations with the help of annotated semantics refinement. Dynamic recognition and determination steps: Introduce a neural network with the ability of time series encoding. Considering the dynamic coherence of power operations, use the long short-term memory network structure to capture the operation sequence and time interval. When processing the time series data of power operations, the input gate determines the degree to which new feature information at the current moment is incorporated into the memory unit; the forget gate controls which information in the memory unit of the previous moment needs to be retained or discarded; the output gate determines which information in the memory unit is output for the current moment's output. For a series of image and audio features corresponding to power operation actions, as time progresses, the long short-term memory network can continuously update the "memory" of the operation process, accurately capture the logical relationship between previous and subsequent operations, and track the features that evolve over time after multi-level injection and enhancement.In the subsequent introduction of a fully connected layer to reduce the dimension of high-dimensional time-series features and condense key information, the classifier finally accurately determines whether the operation is illegal based on the condensed features. Relying on multi-modal fusion, the system can accurately process features in complex lighting and occlusion scenarios, greatly improving the recognition accuracy, reducing the false positive rate, and strengthening the safety defense line of power operations.
[0029] In one embodiment of the present invention, image data of the power operation site is collected from multiple perspectives and at different distances, and environmental and equipment data are obtained, including: collecting visible light images of the power operation site from multiple perspectives and at different distances, collecting heat distribution information of the power operation site through an infrared camera and generating infrared images; recording environmental parameters such as the light intensity, light direction, environmental temperature, and humidity of the power operation site; and collecting real-time operating parameters of on-site electrical equipment.
[0030] Specifically, visible light images of the operation site are collected from multiple perspectives and at different distances. For example, whether the operator wears insulating gloves according to the specifications is clearly visible in the visible light image. The infrared camera is used to collect heat distribution information, with a focus on the thermal imaging of electrical equipment and the operator. Environmental parameters such as light intensity, light direction, environmental temperature, and humidity are recorded. Strong direct light and backlight scenarios will interfere with the quality of visible light images, and the appearance and performance of electrical equipment also vary under different temperatures and humidities. By mastering these environmental data, targeted compensation can be carried out in the subsequent feature adaptation link to better align multi-modal data. The real-time operating parameters of electrical equipment, such as current, voltage, power, etc., are collected and mutually verified with the image data. If a suspected illegal operation is identified in the image, combined with abnormal fluctuations in the equipment operation data, the accuracy and reliability of illegal operation judgment can be improved.
[0031] Among them, the specific steps also include equipment selection and layout planning: According to the characteristics of the power operation scenario, select suitable image acquisition equipment. A visible light camera with high resolution and high frame rate is selected to ensure that the operator's actions and the equipment appearance can be clearly captured; an infrared thermal imager is paired to accurately monitor the temperature distribution of electrical equipment and timely detect potential overheating hazards. Data acquisition synchronization setting: A synchronous acquisition mechanism is constructed to synchronously trigger the visible light and infrared devices, ensuring that different modal data correspond to the same operation moment at the same time. With the help of a high-precision clock synchronization module or network time protocol, the timestamps of each device are strictly calibrated to eliminate the action matching error caused by the time difference and maintain the spatio-temporal consistency of the data, providing accurate materials for subsequent feature fusion. Environmental parameter recording: The environmental data of the operation site are synchronously measured and stored. Light intensity and direction will interfere with the quality of visible light imaging, and temperature and humidity changes may affect the equipment performance and infrared imaging effect. Light sensors and temperature and humidity meters are used to collect the corresponding parameters in real time and associate them with the image data one by one to assist in correcting the feature deviation caused by environmental differences in the subsequent feature adaptation link.
[0032] In an embodiment of the present invention, the data obtained at the power operation site is paired according to the time stamp for different modality data captured at the same moment, including: summarizing the collected visible light images and infrared images, pairing the different modality images captured at the same moment according to the time stamp to form multiple groups of image pairs; performing normalization processing on the paired data to unify the image size and resolution, and performing feature annotation.
[0033] Specifically, this step aims to extract features from the collected data to construct multi-modal data. Summarize the collected visible light images and infrared images, and pair the different modality images captured at the same moment according to the time stamp to form multiple groups of image pairs. For example, at a certain moment, the visible light image in front of the power distribution cabinet is closely related to the corresponding infrared thermal imaging image, locking the same operation moment. Regularize the paired data groups, unify the image size and resolution, and reduce the interference of image specification differences on subsequent feature extraction. Arrange professional power operation and maintenance personnel to observe on site or review the collected images, and label the collected data accurately; according to the key elements of power operation, such as equipment outline and personnel posture, conduct preliminary feature annotation on the visible light and infrared images respectively, and store the annotation information in the corresponding data label set. Provide semantic guidance for subsequent feature adaptation, and at the same time sort out the key feature context of multi-modal data to complete the transformation from raw data to analyzable multi-modal data. Detailedly annotate the operation type and whether there is a violation, and accurately record the violation links for illegal operations, such as illegal switching on and operation without wearing insulating gloves, etc., to provide reliable supervision information for subsequent model training.
[0034] Among them, constructing multi-modal data uses a unified format to standardize the data and map it to a two-dimensional space. The data formats collected by different devices are often different and need to be converted into a common format. For example, image data is unified into JPEG or PNG format, and audio data is converted into standard audio formats such as WAV to eliminate compatibility problems caused by format differences. Further process the image data, such as performing spatial registration on visible light and infrared images. With the help of feature point detection algorithms, such as Scale-Invariant Feature Transform (SIFT) and Speeded-Up Robust Features (SURF), locate the corresponding feature points in different modality images, and then calculate through the transformation matrix to align the images in space to ensure that the positions of the same object in each modality are accurately corresponding. Normalize the numerical range of each modality data, and map pixel values, signal intensities, etc. to a specific interval, such as [0, 1] or [-1, 1]. For image data, through linear transformation and Z-score normalization, balance the magnitude differences of different modality data to avoid a certain modality data dominating in subsequent fusion.
[0035] In one embodiment of the present invention, a deep convolutional neural network is used to align the visual semantics of objects, scenes, and actions in image data with the semantics of words and sentences in language, and combined with environmental and device data to form feature data, including: multi-scale feature extraction, using a deep convolutional neural network to extract feature data of different scales in visible light images and infrared images; feature distribution calibration, using an adaptation module to adjust feature data of different modalities and different scales by means of normalization and linear transformation methods, so that they tend to be consistent in data range and distribution form; semantic space alignment, using feature annotation information and a pre-trained semantic model to map the feature data in visible light images and infrared images to the same semantic space, so as to align the visual semantics of objects, scenes, and actions in the image data with the semantics of words and sentences in the language; feedback optimization, inputting the feature data into a small discriminant model, and reversely fine-tuning the adaptation parameters according to the discriminant results, and iterating repeatedly to improve the fusion of different modality features.
[0036] Specifically, this step aims to eliminate the differences between multi-modal data and achieve accurate semantic alignment. Multi-scale feature extraction stage: Using visible light and infrared images, a deep convolutional neural network is used to extract features of different scales. Feature distribution calibration stage: The power operation scene is easily affected by light, and the infrared is interfered by temperature, resulting in large differences in feature distribution. The adaptation module uses means such as normalization and linear transformation to adjust the mean and variance of features of different modalities and different scales, so that they tend to be consistent in numerical range and distribution form. Semantic space alignment stage: There are standard and normative actions in power operations. Using annotation information and a pre-trained semantic model, the visible light and infrared features are mapped to the same semantic space, so that the semantics are coherent, avoiding recognition deviations caused by modal characteristics, and improving the generality and stability of the overall features. Feedback optimization stage: The features that have been initially adapted are input into a small discriminant model, and the adaptation parameters are reversely fine-tuned according to the discriminant results, and iterated repeatedly until the fusion effect of different modality features reaches the optimal, and high-quality aligned features are delivered to the multi-level feature pruning and injection link.
[0037] Among them, align the features of different modalities and construct an adaptive module to adapt the features. Specifically, it is divided into two steps: Feature Vector Preparation and Similarity Calculation: Feature vectors corresponding to image data of different modalities are intercepted at the same spatial position. For a specific power violation operation frame, a feature vector representing the hand movement contour is extracted at a certain coordinate point in the visible light feature map, and a feature vector reflecting the hand temperature distribution is extracted at the corresponding coordinate in the infrared feature map. The similarity calculation uses dot product operation, cosine similarity, or other similarity metric functions to calculate the correlation score between feature vectors of different modalities. Taking the dot product as an example, the visible light hand movement feature vector and the infrared hand temperature feature vector are multiplied element by element and then summed to obtain a scalar value. The higher this value, the stronger the correlation between these two cross-modal features at this position. Traverse each position of the entire feature map and repeat the above calculation process to construct a correlation score matrix with the same size as the original feature map. Each element of the matrix corresponds to the degree of association of cross-modal features at a specific position in the image.
[0038] Attention Map Generation and Weighted Fusion: The Softmax function is used to convert the correlation score into weights. Each element in the matrix is mapped to a value between 0 and 1, and the sum of elements in each row is 1, thereby generating an attention map (Attention Graph). In essence, it is a weight matrix that highlights the key regions with high cross-modal correlation. Using the generated attention map, weighted summation is performed on the original features of different modalities. For a certain feature channel of the visible light image, each feature value on this channel is multiplied by the weight at the corresponding position in the attention map, and the same operation is performed on the same feature channel of the infrared image. Then, the weighted visible light features and infrared features are added element by element to obtain the fused features. Traverse all feature channels to complete the weighted fusion of the entire feature map, enabling the fused features to adaptively combine the advantageous information of the two modalities in different regions, and strengthening the features closely related to violation operations such as key actions and temperature changes.
[0039] In an embodiment of the present invention, a deep convolutional neural network is used to extract feature data of different scales in visible light images and infrared images, including: fine texture features of device edges and personnel contours are obtained at the bottom layer; geometric shape features related to action postures are captured at the middle layer; and abstract concept features of action semantics and device operating states are focused on at the high layer.
[0040] Specifically, through the layer-by-layer extraction of bottom, middle, and high layer features, a richer feature representation can be formed, covering multi-dimensional information from details to the whole, which can more comprehensively understand the complex scenarios and their dynamic changes in the power operation site.
[0041] In an embodiment of the present invention, multi-layer deep fusion and pruning optimization are performed on the feature data to screen out key features, including: hierarchically dividing the feature data into a bottom layer, a middle layer, and a top layer; injecting operations from the bottom layer to the middle layer, and by means of a feature importance evaluation mechanism, analyzing the contribution degree of the bottom layer features to violation recognition, screening out key bottom layer features, and integrating them into the middle layer in a weighted fusion manner. The features after middle layer fusion are then injected into the top layer; according to the correlation between the middle layer and top layer features, an adaptive fusion rule is constructed to dynamically adjust the fusion coefficient, so as to make the concrete action features in the middle layer match the abstract semantics in the top layer precisely.
[0042] Specifically, it aims to perform deep fusion and pruning optimization on calibrated features at different levels to strengthen the key features closely related to power violation operations. First, the features are hierarchically divided into a bottom layer, a middle layer, and a top layer. The injection operation goes from the bottom layer to the middle layer. By means of a feature importance evaluation mechanism, the contribution degree of the bottom layer features to violation recognition is analyzed, key bottom layer features are screened out, and they are integrated into the middle layer in a weighted fusion manner. The features after middle layer fusion are then injected into the top layer. According to the correlation between the middle layer and top layer features, an adaptive fusion rule is constructed to dynamically adjust the fusion coefficient, so as to make the concrete action features in the middle layer match the abstract semantics in the top layer precisely. The limb actions corresponding to illegal switching are integrated into the top layer illegal semantics to achieve deep coupling of semantics and actions. During this process, small-scale test data is used to monitor the effect, and parameters are fine-tuned as needed to ensure that violation features are prominent. At the same time, pruning operations are performed before injecting features at each level to lighten the model.
[0043] Among them, the interaction of hierarchical features is used to improve the fusion effect of multi-modal information, and at the same time, a pruning strategy is introduced to reduce the number of parameters and the amount of computation. This step is divided into four steps: Pruning strategy formulation: First step, calculate the L1 norm of each channel feature map, which intuitively reflects the overall intensity of the signal within a channel. Second step, sort according to the L1 norm, and select the channels to be retained according to a percentage. Third step, remove the corresponding channels, adjust the corresponding dimensions, and fine-tune the model. Use a subset of the original training data to retrain the model with a smaller learning rate, so that the model parameters adapt to the new simplified structure and recover the accuracy loss. The above pruning process is repeated before the feature injection link at each level.
[0044] Underlying injection: After being optimized by the feature adaptive selection module, the underlying features of visible light and infrared images are obtained. These underlying features retain high-resolution information and contain fine textures such as device edges and initial movements of personnel's hands. To ensure that the two can be fused in the same computational framework, it is necessary to unify the size of the feature matrices so that the underlying features of visible light and infrared strictly correspond in the spatial dimension. Considering the characteristics of the power scenario, weights are assigned to the underlying features of visible light and infrared. The visible light image captures actions precisely, and a higher weight is given to highlight action details; the infrared image reflects temperature information, and for device operations that may involve abnormal heating, the weight of its temperature-related features is appropriately increased. The weighted average formula is used: , where is The fused underlying features, 、 are weights, and are the underlying features of visible light and infrared respectively. The fused underlying features are incorporated into the middle-level feature processing process according to the channel dimension or feature splicing method.
[0045] Middle-level injection: With the dynamic weight strategy as the core, the weights are adjusted according to the real-time scene complexity of power operations. In simple operation scenarios, the weights of the underlying supplementary features are slightly lower; in complex maintenance process scenarios, the weights of the precise starting actions brought by the underlying layer increase, making the fusion more focused on key local dynamics. Using convolutional operations, feature extraction and enhancement are performed on the fused middle-level features. Focusing on local structural information such as the coherence of limb movements and tool operation trajectories in power operations, the enhanced middle-level features are injected into the high-level features.
[0046] High-level injection: When the high-level features receive the middle-level features, through means such as feature splicing and weighted fusion, a close connection is established between the action postures in the middle level and the high-level semantics. For example, the action posture features of "illegal switch closing" extracted in the middle level are aligned and fused with the existing operation semantic features in the high level in the feature space. The fused high-level features are retrained and optimized using the rich violation semantic information in the labeled dataset. The high-level feature representation is adjusted with the feedback of the classifier.
[0047] In an embodiment of the present invention, based on a neural network with the ability of time series coding, key features input to the neural network with the ability of time series coding in chronological order are dynamically recognized, and a determination result of whether the power operation is an illegal operation is output, including: using a neural network with the ability of time series coding as the basic framework to process the time series features of power operation behaviors and retain key information; inputting the key features into the neural network in chronological order so that the neural network learns the evolution law of the action states at different time points; connecting a classifier, and outputting a determination result of whether the power operation is an illegal operation based on the integrated features.
[0048] Specifically, it aims to utilize the features after pre - stage fusion, adaptation, and injection processing to build a system architecture that can accurately capture the dynamic process of power operations and determine whether there are violations. First, a neural network with the ability of temporal encoding is selected as the basic framework to process the characteristics of operation behaviors in power operations that unfold over time series, and remember the key feature information left by previous operation steps. Second, the multi - level feature pruning, injection, and highly condensed feature data are sequentially input into this neural network in chronological order, enabling the network to gradually learn the evolution law of action states at different time points. Finally, a SoftMax classifier is connected, and the final discrimination result is output based on the integrated features to ensure the reliability and accuracy of power operation recognition.
[0049] In an embodiment of the present invention, processing the temporal features of power operation behaviors and retaining key information includes: using a fully - connected layer to perform dimensionality reduction on high - dimensional temporal feature vectors, integrating and refining key information, and reorganizing the scattered features according to the violation discrimination logic.
[0050] Specifically, using a neural network with the ability of temporal encoding to capture the dynamic operation process, and then through the processing of the fully - connected layer and the classifier, an accurate recognition result of violation operations is output, which is carried out in three steps: Temporal encoding integration: Introduce a neural network with the ability of temporal encoding, such as Long Short - Term Memory (LSTM), Gated Recurrent Unit (GRU), etc. Such recurrent neural network structures can capture the sequential order, time intervals, etc. of operation steps, connect the key features scattered at different time points, form a complete operation pipeline, and provide a dynamic perspective for accurately identifying violation operations.
[0051] Fully - connected layer connection and dimensionality reduction: Introduce a fully - connected layer to perform dimensionality reduction on high - dimensional temporal feature vectors, integrate and refine key information, and reorganize the scattered features according to the violation discrimination logic, enabling the subsequent classifier to carry out discrimination work in a more compact and key feature space, reducing the interference of redundant information.
[0052] Classifier accurate discrimination: Adopt mature classification algorithms such as Support Vector Machine (SVM), Multilayer Perceptron, etc. According to the highly condensed features output by the fully - connected layer, classify whether the operation belongs to the category of violation or compliance. At the same time, during the training process, the classifier continuously optimizes the discrimination boundary, adapts to complex and changeable power operation scenarios, minimizes the misjudgment probability, and outputs a reliable violation recognition result.
[0053] In an embodiment of the present invention, the neural network with the ability of temporal encoding includes Long Short - Term Memory (LSTM).
[0054] Specifically, the long short-term memory network includes an forgetting gate, an input gate, and an output gate, which can dynamically select to retain or forget information. This mechanism enables the network to automatically adjust the degree of attention to features at different time steps, providing higher flexibility for feature selection and avoiding the vanishing gradient problem in long-term dependence learning of traditional neural networks.
[0055] In an embodiment of the present invention, the acquisition frequency of power operation data is adjusted according to the complexity of power operations. The higher the complexity of power operations, the higher the acquisition frame rate, so as to ensure continuous capture of actions.
[0056] Specifically, by dynamically adjusting the acquisition frequency according to the complexity of power operations, each detail of key operations can be captured more accurately. Especially in complex power operations, such as equipment maintenance, high-voltage operations, or fault troubleshooting, etc., it can ensure that every important action is continuously recorded. Specific embodiments: Such as Figure 2 shown, specifically including the following steps: Power operation data acquisition step: Fully consider the complexity and particularity of the power operation scenario, select suitable image acquisition devices, infrared thermal imagers, and collect visible light image data, infrared image data, environmental data, and equipment operation data. At the same time, build a synchronous acquisition mechanism, and with the help of a high-precision clock synchronization module or network time protocol, ensure that different modality data at the same moment accurately corresponds to the same operation instant, maintaining spatio-temporal consistency. In addition, use a light sensor and a temperature and humidity meter to collect the environmental data of the operation site in real time, store it in an associated manner, assist in correcting feature deviations subsequently, and also arrange professional operation and maintenance personnel to label operation tags, specifying the operation type and whether there are violations in detail, providing reliable supervision information for model training.
[0058] Steps for constructing multi-modal data: The initial formats of the data collected by various devices are different and need to be converted into a common format for easy processing. Both visible light and infrared images are converted into the JPEG format, and audio-type environmental monitoring data is converted into the WAV format. The audio with different sampling frequencies is unified into the standard sampling frequency required by the model; the audio signal is frame-processed, and the continuous audio is cut into short frame segments of a fixed duration to prepare for feature extraction.
[0059] For visible light and infrared images, use the SIFT feature point detection algorithm to search for corresponding feature points among the scale-invariant feature points detected from different modality images. Suppose the two images are and , and the sets of corresponding feature points detected are and , calculate the transformation matrix through optimization algorithms such as the least squares method, so that 。Since there are significant differences in the data ranges of each modality, it is necessary to normalize the data. For image pixel values, linear transformation is adopted. Let the original pixel value be , and the transformation formula is: The above formula scales the eigenvalues to a small specific interval, usually [0, 1]; reducing the impact of the dimensional differences between different features on model training. Among them is the normalized eigenvalue.
[0060] For audio data, Mel-Frequency Cepstral Coefficients (MFCC) are extracted.
[0061] First step, boost the high-frequency components through pre-emphasis: Among them is the original audio, is the audio after pre-emphasis, and the coefficient is usually taken as 0.97.
[0062] Second step, frame and window the continuous audio: Assume the audio signal , the frame length is , the frame shift is , then the -th frame signal can be expressed as: After framing, a Hamming window is used for windowing: The windowed frame signal is: .
[0063] Third step, introduce the Fast Fourier Transform (FFT): Perform FFT on the windowed signal of each frame to convert the time-domain signal to the frequency domain: Among them is the discrete Fourier transform result of the -th frame signal, which represents the amplitude and phase information of the frame signal at different frequencies.
[0064] Fourth step, calculate the Mel spectrum: The conversion formula between Mel frequency and linear frequency is: Convert the FFT result to the Mel frequency scale. First, define the Mel filter bank. Suppose there are Mel filters , and the Mel spectrogram is calculated as follows: Here is the square of the FFT magnitude spectrum, and the Mel filter bank is triangular, covering different Mel frequency ranges and used for weighted summation of the spectral energy. Take the natural logarithm of the Mel spectrogram to obtain the log Mel spectrogram : Finally, perform the discrete cosine transform (DCT) to extract the Mel frequency cepstral coefficients : where is the number of MFCC coefficients to be retained, usually taking 12 - 13, which contains the characteristic information of the signal.
[0065] Feature adaptation step: This step aligns the features of different modalities and constructs an adaptive module to adapt the features. Multi-scale feature extraction: Suppose the multi-modal data includes visible light image data and infrared image data . Input them into a deep convolutional neural network (Convolutional Neural Network, CNN). The visible light feature map output by the l th layer of the CNN is , and the infrared feature map is . For a given position , intercept the feature vectors and from the feature map. Calculate the correlation score between the feature vectors using the cosine similarity. The formula is: where is the dimension of the feature vector; traverse the entire region of the feature map with size to generate the correlation score matrix . Then, use the Softmax function to generate the attention map. The formula is as follows: where it is required that . After generating the attention map, perform weighted fusion. Suppose a certain feature channel of the visible light feature map is , and that of the infrared is . The fused feature channel The calculation formula is as follows: Repeat this operation for all feature channels to complete the fusion; output the adaptive features rich in violation discriminative power.
[0066] Feature pruning and injection steps: Use the interaction and pruning optimization of hierarchical features to improve the fusion effect of multi-modal information. Before feature injection and fusion at each level, a pruning operation based on the L1 norm is required. For the features obtained by convolution (where represents the number of channels, and represent the height and width of the feature map respectively), for each channel , calculate its L1 norm, and the formula is: where is the element value of the feature at the channel, with height and width . The larger the L1 norm, the more effective information the channel contains. Sort these channels in descending order according to the value of the L1 norm. Set the pruning threshold according to the proportion of retained channels. Traverse all channels, and mark the channels with L1 norm less than as the channels to be pruned. Remove these marked channels from the feature to obtain the pruned feature .
[0067] Subsequently, perform the feature injection operation: At the bottom layer, let the visible light bottom layer feature processed by the feature adaptive selection module be , with dimension , and the infrared bottom layer feature be , with dimension . First, unify the sizes through methods such as interpolation to make , and obtain and . Considering the power scenario, let the weight given to the visible light bottom layer feature be , and the weight of the infrared bottom layer feature be , and . For weighted average fusion, the calculation formula of the fused bottom layer feature is: Let the middle layer feature be . Incorporate the fused bottom layer feature into the middle layer by concatenating along the channel dimension. The concatenated middle layer feature can be expressed as: , where concat represents the concatenation operation along the channel dimension. Dynamic weight adjustment is performed during the middle layer injection stage. Let the weight of the features injected at the bottom layer during middle layer fusion be , and the original feature weight of the middle layer be , the dynamic weight is adaptively adjusted according to the scene complexity and can be simply set as , where is the adjustment coefficient, sigmoid and the function is used to map the value to the
[0068] Convolutional enhancement and injection of high-level features: Perform a convolutional operation on the fused middle layer features , and let the convolutional kernel be . The enhanced middle layer features . Inject it into the high-level features , and obtain the new high-level features The high-level injection action posture feature weight , the semantic feature weight , and the fused high-level features are: Finally, semantic optimization is performed. Let the fused high-level features be , and the corresponding semantic label be . Use the cross-entropy loss function for retraining, where is the number of samples, is the number of classes, is the true label, is the predicted probability. According to the feedback of the classifier, backpropagation is used to update the high-level feature representation to make the high-level features accurately associated with the violation semantics.
[0069] Dynamic recognition and determination steps: Let the feature sequence after multi-level feature pruning and injection be , where is the time step, is the feature vector at time , and the dimension is . Taking the long short-term memory network (LSTM) as an example, its calculation process is as follows: Input gate: Forget gate: Candidate memory unit: Memory unit: Output gate: Hidden state: Where represents the weight matrix, represents the bias vector, is sigmoid function, is the element-wise multiplication. The hidden state sequence integrating the temporal information is obtained through the LSTM layer . Let the weight matrix of the fully connected layer be , with the dimension of , and the bias vector be . The calculation formula is: After compressing the high-dimensional temporal features to low dimensions, it is input into a multi-layer perceptron (MLP). Let the hidden layer weight matrix be , the bias be , the output layer weight matrix be , and the bias be . The output of the hidden layer is , and the prediction result of the output layer is . softmax The function converts the output into a categorical probability distribution, and the category with the highest probability is taken as the prediction result to accurately determine whether the operation is illegal.
[0070] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.
Claims
1. A method for identifying electricity violation operations based on multimodal fusion, characterized in that, Including: Power operation data collection, collecting image data of the power operation site from multiple perspectives and different distances, and obtaining environmental and equipment data; Constructing multimodal data, pairing the data obtained from the power operation site according to the timestamp, pairing different modal data captured at the same moment; Feature adaptation, using a deep convolutional neural network to align the visual semantics of objects, scenes, and actions in the image data with the semantics of words and sentences in the language, and combining the environmental and equipment data to form feature data; Feature pruning and injection, performing multi-layer deep fusion and pruning optimization on the feature data to screen out key features; Dynamic recognition and determination, based on a neural network with temporal coding ability, dynamically recognizing the key features input to the neural network with temporal coding ability in the order of timestamps, and outputting a determination result on whether there is an illegal operation in the power operation.
2. The method for identifying power violation operations based on multimodal fusion according to claim 1, wherein The collecting the image data of the power operation site from multiple perspectives and different distances, and obtaining environmental and equipment data includes: Collecting visible light images of the power operation site from multiple perspectives and different distances, collecting heat distribution information of the power operation site through an infrared camera and generating infrared images; recording environmental parameters such as the light intensity, light direction, environmental temperature, and humidity of the power operation site; and collecting real-time operating parameters of on-site electrical equipment.
3. The method for identifying power violation operations based on multi-modal fusion according to claim 2, characterized in that The pairing the data obtained from the power operation site according to the timestamp, pairing different modal data captured at the same moment includes: Summarizing the collected visible light images and infrared images, pairing different modal images captured at the same moment according to the timestamp to form multiple groups of image pairs; performing standardization processing on the paired data, unifying the image size and resolution, and performing feature annotation.
4. The method for identifying power violation operations based on multi-modal fusion according to claim 3, characterized in that, The using a deep convolutional neural network to align the visual semantics of objects, scenes, and actions in the image data with the semantics of words and sentences in the language, and combining the environmental and equipment data to form feature data includes: Multi-scale feature extraction, using a deep convolutional neural network to extract feature data of different scales in the visible light image and the infrared image; Feature distribution calibration, using an adaptation module to adjust feature data of different modalities and different scales by using normalization and linear transformation methods to make them tend to be consistent in the data range and distribution form; Semantic space alignment, using feature annotation information and a pre-trained semantic model to map the feature data in the visible light image and the infrared image to the same semantic space, so that the visual semantics of objects, scenes, and actions in the image data are aligned with the semantics of words and sentences in the language; Feedback optimization, inputting the feature data into a small discriminant model, and reversely fine-tuning the adaptation parameters according to the discriminant result, and iterating repeatedly to improve the fusion of different modal features.
5. The method for identifying power violation operations based on multimodal fusion according to claim 4, characterized in that, The using a deep convolutional neural network to extract feature data of different scales in the visible light image and the infrared image includes: Obtaining fine texture features of device edges and personnel contours at the bottom layer; capturing geometric shape features related to action postures at the middle layer; focusing on abstract concept features of action semantics and device operating states at the high layer.
6. The method for identifying power violation operations based on multimodal fusion according to claim 1, wherein Performing multi-layer deep fusion and pruning optimization on the feature data to screen out key features, including: Dividing the feature data into a bottom layer, a middle layer, and a top layer hierarchically; injecting operations from the bottom layer to the middle layer, analyzing the contribution of bottom-layer features to violation recognition with the help of a feature importance evaluation mechanism, screening out key bottom-layer features, and integrating them into the middle layer using a weighted fusion method. The features after middle-layer fusion are then injected into the top layer; based on the correlation between middle-layer and top-layer features, constructing an adaptive fusion rule, dynamically adjusting the fusion coefficient, and promoting the precise matching of the concrete action features in the middle layer with the abstract semantics in the top layer.
7. The method for identifying electricity violation operations based on multimodal fusion according to claim 1, characterized in that, Based on a neural network with time series encoding ability, dynamically identifying the key features input to the neural network with time series encoding ability in the order of time stamps and outputting a determination result on whether the power operation is an illegal operation, including: Using a neural network with time series encoding ability as the basic framework to process the time series features of power operation behaviors and retain key information; Inputting the key features into the neural network in the order of time stamps, enabling the neural network to learn the evolution law of the action states at different time points; Connecting a classifier and outputting a determination result on whether the power operation is an illegal operation based on the integrated features.
8. The method for identifying power violation operations based on multimodal fusion according to claim 7, characterized in that, Processing the time series features of power operation behaviors and retaining key information, including: Using a fully connected layer to perform dimensionality reduction on the high-dimensional time series feature vector, integrating and refining key information, and reorganizing the scattered features according to the violation discrimination logic.
9. The method for identifying illegal power operation based on multi-modal fusion according to claim 1, characterized in that, The neural network with time series encoding ability includes a long short-term memory network.
10. The method for identifying electricity violation operations based on multi-modal fusion according to any one of claims 1 to 9, characterized in that, The acquisition frequency of the power operation data acquisition is adjusted according to the complexity of the power operation. The higher the complexity of the power operation, the higher the acquisition frame rate, so as to ensure continuous capture of actions.
Citation Information
Cited By
Intelligent monitoring and identifying system for violation of power transmission line construction site
CN120564366A
Video analysis-based multi-scene operator violation behavior identification method and system
CN120726699A
Method, device, system and equipment for detecting violation state of individual protection equipment
CN121033770A
Methods, devices, systems and equipment for detecting violations of personal protective equipment
CN121033770B
Power station operator violation behavior identification method and device based on multi-modal fusion
CN122416350A