A lightweight image recognition method and system for power plant safety

By constructing a multi-scale adaptive perception and binocular stereo mapping model, combined with a lightweight knowledge distillation model, the problem of identifying violations by power plant workers was solved, enabling real-time monitoring and early warning of power plant safety.

CN120766183BActive Publication Date: 2025-11-25Beijing Huadian Wanfang Certification Co., Ltd.
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510891960.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-25
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing technologies lack a deep understanding of the complex behavioral patterns and spatiotemporal evolution of power plants, making it difficult to identify violations by operators, leading to escalation of safety hazards and accidents.

Method used

A multi-scale adaptive perception image anomaly detection model and a multi-stage binocular stereo mapping model are constructed, combined with a lightweight knowledge distillation edge device inference model, to monitor power plant safety hazards in real time.

Benefits of technology

It enhances the ability to identify abnormal behaviors in power plant operation scenarios, achieves efficient safety early warning and safety distance calculation, and reduces computing latency and memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766183B_ABST
    Figure CN120766183B_ABST
Patent Text Reader

Abstract

The application discloses a kind of light-weight image recognition method and system for power plant safety, it is related to image recognition field, the method comprises: constructing multiscale adaptive perception's image anomaly detection model, utilizes image anomaly detection model to identify the illegal behavior of power plant operating personnel in power plant real-time video data;Multi-stage binocular stereo area matching model is constructed to the feature point in the binocular image of electric power equipment is matched, and the safety distance of power plant operating personnel and live body is calculated based on the matching result;Combining the illegal behavior of power plant operating personnel and safety distance, construct light-weight knowledge distillation's edge device reasoning model, and utilize edge device reasoning model real-time monitoring power plant safety hidden danger.The application will have high expression ability's teacher model be deployed in cloud training, and pass semantic information and structure knowledge to student model on edge device by dynamic adaptation strategy, significantly improve the recognition ability of edge model in limited scene of computing power.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image recognition, in particular to a lightweight image recognition method and system for power plant safety. BACKGROUND

[0002] Power plant safety refers to ensuring the overall safety of personnel, equipment and systems during power production, transmission and operation, covering multiple aspects such as operation specification, explosion prevention and touch prevention, electrical isolation, warning system, etc. Lightweight image recognition refers to deploying an image recognition model with fast reasoning, low power consumption and high adaptability on edge devices in a high-risk environment of a power plant to ensure the safety of operating personnel and equipment operation, and realizing efficient identification and real-time warning of typical violation behaviors such as safety helmet wearing, insulating glove use, overline behavior and intrusion into live area.

[0003] The existing image recognition technology lacks deep understanding and dynamic adaptation capability for complex behavior patterns and spatio-temporal state evolution, making it difficult to identify high-risk behaviors such as operating personnel not wearing insulating gloves, crossing safety distance line and violating into live area, resulting in insufficient abnormal identification capability under sudden behavior changes or shielding interference, and further causing hidden trouble escalation and even accidents.

[0004] At present, there is no effective solution to the problems in the related art. SUMMARY

[0005] In view of the problems in the related art, the present application proposes a lightweight image recognition method and system for power plant safety to overcome the above technical problems existing in the prior art.

[0006] To this end, the specific technical solutions adopted by the present application are as follows:

[0007] According to one aspect of the present application, a lightweight image recognition method for power plant safety is provided, which comprises:

[0008] acquiring real-time video data of the power plant, constructing a multi-scale adaptive perception image anomaly detection model, and identifying the violation behavior of the power plant operating personnel in the real-time video data of the power plant by using the image anomaly detection model;

[0009] acquiring binocular images of power equipment, constructing a multi-stage binocular stereo region matching model to match feature points in the binocular images of power equipment, and calculating the safety distance between the power plant operating personnel and the live body based on the matching result;

[0010] combining the violation behavior of the power plant operating personnel and the safety distance, constructing a lightweight knowledge distillation edge device reasoning model, and using the edge device reasoning model to monitor the safety hidden danger of the power plant in real time.

[0011] Preferably, the multi-scale adaptive perception image anomaly detection model is constructed, and the power plant operation personnel's illegal behavior in the power plant real-time video data is identified by using the image anomaly detection model, which includes:

[0012] The visual features and spatial topological features in the power plant real-time video data are mined by using a bilinear convolutional neural network, and the visual features and spatial topological features are fused to identify the reference power operation scene;

[0013] Based on the reference power operation scene, the context information is encoded by using a time sequence context detector in the transition perception context network, and the power plant real-time video data is segmented by using a transition perception classifier to form behavior video clips and transition video clips;

[0014] The multi-scale feature pyramid network and the adaptive attention mechanism are combined to identify the operation abnormal behavior of the power plant operation personnel in the transition video clips;

[0015] The time series model is used to identify the posture abnormal behavior of the power plant operation personnel in the behavior video clips.

[0016] Preferably, the multi-scale feature pyramid network and the adaptive attention mechanism are combined to identify the operation abnormal behavior of the power plant operation personnel in the transition video clips, which includes:

[0017] The transition image sequence is extracted from the transition video clips, and the multi-scale feature pyramid network is constructed based on the transition image sequence;

[0018] The sparse optical flow field of the transition image in the topmost layer of the multi-scale feature pyramid network is calculated, and the inter-frame difference processing is performed on the adjacent transition images to obtain the initial optical flow vector of the adjacent transition images;

[0019] The optical flow vectors of the remaining transition images are calculated based on the initial optical flow vector, and are fused with the initial optical flow vector to obtain the fused optical flow vector;

[0020] The fused optical flow vector is compared with a preset threshold value, if the fused optical flow vector is greater than the preset threshold value, then the optical flow vector of the maximum transition image is selected as an abnormal point and is modified to the mean value of the optical flow vector, otherwise, the fused optical flow vector is transmitted to the next layer until it is transmitted to the base layer of the multi-scale feature pyramid network;

[0021] The fused optical flow vector of each layer is converted into a depth map, the depth map is channel spliced based on the attention mechanism, and the operation abnormal behavior is segmented from the spliced channel by using the transition perception context network.

[0022] Preferably, the power equipment binocular image is obtained, the multi-stage binocular stereo region matching model is constructed to match the feature points in the power equipment binocular image, and the safety distance between the power plant operation personnel and the live body is calculated based on the matching result, which includes:

[0023] acquire binocular images of the power equipment by using a binocular camera, and construct a multi-stage binocular stereo area matching model to match the binocular images to obtain a parallax image;

[0024] acquire a depth value of a point of the power equipment in the parallax image by using a triangulation method, and reconstruct a three-dimensional coordinate point of the parallax image by using an improved semi-global block matching algorithm;

[0025] based on the abnormal behavior of the posture of the power plant worker, calculate the spatial distance between the reconstructed three-dimensional coordinate point and the feature point of the target object, and determine the safety distance between the power plant worker and the live body according to the spatial distance;

[0026] determine an electronic fence in the live body area according to the safety distance, and monitor whether the power plant worker is in the electronic fence area, and if the power plant worker exceeds the electronic fence area, then prompt the power plant worker to re-enter the electronic fence area through voice.

[0027] Preferably, based on the abnormal behavior of the posture of the power plant worker, calculating the spatial distance between the reconstructed three-dimensional coordinate point and the feature point of the target object, and determining the safety distance between the power plant worker and the live body according to the spatial distance comprises:

[0028] The abnormal behavior of the posture of the power plant worker is decomposed into atomic action units, a picture area of a preset size is intercepted with the moving target as the center point, the picture area is divided into small areas of a preset size, and a plurality of prior anchor frames of different sizes are set at the center points of each small area.

[0029] A power plant worker recognition model is established, and the prior anchor frames of different sizes are distributed into target areas of different scales according to their sizes by the power plant worker recognition model.

[0030] The target class confidence and the target position confidence of each target area are calculated, and the confidence calculation result is taken as the prediction result of the prior anchor frame, and the confidence is corrected according to the existence of the target and the target class in the prior anchor frame.

[0031] Based on the prediction result of the corrected prior anchor frame, the non-maximum suppression algorithm is used to remove the repeated prediction results as the output result of the power plant worker recognition model.

[0032] Based on the output result of the power plant worker recognition model, the feature points of the target object are extracted, the three-dimensional coordinates of the feature points of the target object and the two-point surface centers of the live body are obtained, and the safety distance between the power plant worker and the live body is calculated.

[0033] Preferably, the confidence is corrected according to the existence of the target and the target class in the prior anchor frame comprises:

[0034] The double causal paths are established based on the existence of a target inside a prior anchor frame, a target category and a predefined environmental interference factor, and an intervention variable is introduced to test the double causal paths;

[0035] The mutual information of the double causal paths is calculated by constructing a main branch encoder and an auxiliary branch encoder respectively, and the mutual information is used to perform causal intervention inside the prior anchor frame through the double causal paths to evaluate the influence of the environmental interference factor on the confidence;

[0036] Based on the influence degree of the environmental interference factor on the confidence, the confidence is corrected by introducing interference compensation and category prior probability, and the corrected confidence is injected back into the main branch encoder to make the output of the main branch encoder consistent with the semantics of the corrected confidence.

[0037] Preferably, in combination with the rule-breaking behavior of power plant workers and the safety distance, a lightweight knowledge distillation edge device reasoning model is constructed, and the edge device reasoning model is used to monitor the power plant safety hazards in real time, including:

[0038] The rule-breaking behavior of power plant workers and the safety distance are used to construct a lightweight knowledge distillation edge device reasoning model based on a convolutional network;

[0039] The edge device reasoning model is compressed by operator optimization technology, tensor optimization technology and data optimization technology respectively;

[0040] Based on the model compression result, the power plant electrical safety hazards are identified, and the performance of the edge device reasoning model is optimized by combining self-supervised contrastive learning and meta-learning to adapt to new types of power plant electrical safety hazards;

[0041] The pre-defined teacher model is deployed in the cloud for training, and knowledge is transferred to the lightweight student model through a dynamic adaptation strategy to optimize the reasoning ability of the edge device reasoning model.

[0042] Preferably, the pre-defined teacher model is deployed in the cloud for training, and knowledge is transferred to the lightweight student model through a dynamic adaptation strategy to optimize the reasoning ability of the edge device reasoning model, including:

[0043] The sample set is obtained by integrating real-time video data of the power plant and binocular images of power equipment, and the sample set is divided into labeled data and input into the teacher model deployed in the cloud and the local lightweight student model;

[0044] The prediction labels are output by the teacher model and the local lightweight student model respectively, the prediction samples with clear labels are introduced into the classification loss function, and the label difference between the prediction labels and the true labels of the student model is calculated;

[0045] The label difference based on the student model adjusts the weight of the teacher model, and the knowledge is transferred to the student model based on the optimized teacher model through a dynamic adaptation strategy, and the weight of the teacher model is adjusted reversely;

[0046] The teacher model weight updating and knowledge distillation are iteratively performed until the iteration condition is met, the knowledge between the teacher model and the student model is alternated, and the inference ability of the edge device inference model is optimized.

[0047] Preferably, the formula of the confidence correction is:

[0048]

[0049] In the formula, Pr(Class i ) represents the probability that the sample belongs to the class i; Pr(Class i |Object) represents the possibility of predicting the target data as the i-th class; and Pr(Object) represents the possibility of the existence of the target in the region. represents the intersection over union between the predicted target and the actual target.

[0050] According to another aspect of the present application, a lightweight image recognition system for power plant safety is also provided, which comprises:

[0051] The rule violation recognition module is configured to acquire real-time video data of the power plant, construct a multi-scale adaptive perception image anomaly detection model, and identify the rule violation behavior of the power plant operating personnel in the real-time video data of the power plant by using the image anomaly detection model.

[0052] The safety distance calculation module is configured to acquire binocular images of power equipment, construct a multi-stage binocular stereo region matching model to match feature points in the binocular images of the power equipment, and calculate the safety distance between the power plant operating personnel and the live body based on the matching result.

[0053] The safety hazard monitoring module is configured to combine the rule violation behavior of the power plant operating personnel and the safety distance, construct a lightweight knowledge distillation edge device inference model, and monitor the safety hazards of the power plant in real time by using the edge device inference model.

[0054] The present application has the following advantages:

[0055] 1、The present application effectively enhances the model's perception of the relationship between personnel, objects and background in the work scene by synchronously mining the visual features and spatial topology of the real-time video data in the power plant, introduces a transition perception context network to accurately divide the behavior video segments and behavior transition segments, and uses a multi-scale feature pyramid structure and an adaptive attention mechanism to focus on the subtle action changes in the transition segments, identify fine-grained dangerous actions such as bending over the line and hand contact violations, and further enhance the time sensitivity of the behavior evolution process, meeting the fine identification and early warning needs of abnormal behavior in high-risk, real-time and safety-constrained scenes in the power plant.

[0056] 2、The present application acquires binocular images of power equipment, constructs a multi-stage stereo matching model to accurately match feature points in the images, generates a high-quality disparity map, and realizes spatial reconstruction of the power plant work site by combining an improved semi-global block matching algorithm, and calculates the spatial distance between the abnormal posture behavior of the workers and the live body to scientifically determine the safety distance range between personnel and dangerous sources, thereby effectively enhancing the comprehensive perception ability of work behavior and spatial risk.

[0057] 3、The present application constructs a lightweight knowledge distillation inference model based on a convolutional network, deploys a teacher model with high expression ability in the cloud for training, and transmits semantic information and structural knowledge to a student model on the edge device through a dynamic adaptation strategy, significantly improving the recognition ability of the edge model in the limited computing power scene; at the same time, the model is cut at the structure level and data level by combining compression technologies such as operator optimization, tensor rearrangement and data format reconstruction, significantly reducing memory occupation and calculation delay; based on the compressed model, electrical safety hazards are efficiently identified in the power plant work environment, thereby constructing a high-reliability power plant safety intelligent perception system that balances performance, resource efficiency and self-learning ability. BRIEF DESCRIPTION OF DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0059] Figure 1 is a flowchart of a lightweight image recognition method for power plant safety according to an embodiment of the present application;

[0060] Figure 2 is a principle block diagram of a lightweight image recognition system for power plant safety according to an embodiment of the present application;

[0061] Figure 3is a security protection measure identification technical framework diagram based on fine granularity detection in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0062] Figure 4 is a binocular parallel structure imaging schematic diagram in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0063] Figure 5 is a Census transformation schematic diagram in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0064] Figure 6 is a quadratic curve interpolation schematic diagram in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0065] Figure 7 is a power operation near electric distance detection flowchart based on binocular stereo matching in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0066] Figure 8 is a YOLO algorithm network structure diagram in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0067] Figure 9 is a tensor optimizer schematic diagram in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0068] Figure 10 is a pre-quantization Relu operation schematic diagram in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0069] Figure 11 is a post-quantization Relu operation schematic diagram in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application;

[0070] Figure 12 is a convolutional neural network parameter quantization flowchart in a power plant safety-oriented lightweight image recognition method according to an embodiment of the application.

[0071] In the figure:

[0072] 1, violation behavior identification module; 2, safety distance calculation module; 3, safety hazard monitoring module. DETAILED DESCRIPTION

[0073] To further illustrate the embodiments, the present application provides drawings which are part of the disclosure of the present application, mainly used to illustrate the embodiments, and can be used to explain the operation principle of the embodiments in conjunction with the related description of the specification. Those of ordinary skill in the art can understand other possible implementations and advantages of the present application in conjunction with these contents. The components in the drawings are not drawn to scale, and similar component symbols are generally used to represent similar components.

[0074] According to an embodiment of the present application, a power plant safety-oriented lightweight image recognition method and system are provided.

[0075] The present application will be further described in conjunction with the drawings and specific embodiments. As shown in the drawings and specific embodiments, according to the power plant safety-oriented lightweight image recognition method of the embodiment of the present application, the method comprises: Figure 1

[0076] S1, acquiring real-time video data of the power plant, constructing a multi-scale adaptive perception image anomaly detection model, and using the image anomaly detection model to identify the violation behavior of the power plant operating personnel in the real-time video data of the power plant.

[0077] Among them, constructing a multi-scale adaptive perception image anomaly detection model, and using the image anomaly detection model to identify the violation behavior of the power plant operating personnel in the real-time video data of the power plant comprises:

[0078] Using a bilinear convolutional neural network to mine visual features and spatial topological features in the real-time video data of the power plant, and fusing the visual features and spatial topological features to identify the reference power operation scene.

[0079] It should be noted that the intelligent identification and early warning technology based on AI image recognition technology is based on the violation behavior of the power plant operation site. By using a lightweight mobile edge terminal with AI computing capability, using target identification and detection technology, fine intelligent identification and real-time monitoring of behaviors such as not wearing a safety helmet, not wearing a safety belt correctly, improper dress, smoking, etc. are realized, and voice alarm is performed according to the identification result.

[0080] The purpose of power plant near-electric operation scene recognition is to fully mine the visual features in the image and the spatial topological relationship features between production safety objects according to the perceived real-time video data of the power plant, realize the cognition and understanding of the operation scene and type, and lay a foundation for safety risk identification.

[0081] Considering the special problems in power plant near-electric operation scene understanding, the present application adopts a power scene understanding framework based on cross-modal fusion of visual features and spatial topological features.

[0082] ​Based on the reference power operation scene, the context information is encoded by using a timing context detector in the transition perception context network, and the power plant real-time video data is segmented by using a transition perception classifier to form behavior video segments and transition video segments.

[0083] It should be noted that, in view of the fine identification requirement of personnel and other objects for power plant production safety control, the power production safety hidden danger object identification technology based on fine-grained target detection is proposed, considering that the feature difference between power plant production safety objects is small, and therefore, the fine-grained detection method based on bilinear convolutional neural network is proposed to improve the identification performance of power production safety hidden danger objects, and the specific architecture of the method is as shown in Figure 3 .

[0084] It can be known from Figure 3 that the fine-grained target detection of the bilinear convolutional neural network is a four-tuple, specifically:

[0085] F=(f A ,f B ,P,C);

[0086] In the formula, F represents the fine-grained target detection result; f A and f B represent the feature extraction functions of two bilinear convolutional neural networks A and B in Figure 3 ; P represents a pooling function; and C represents a classification function for classifying and identifying the bilinear feature vector after a Softmax normalization layer.

[0087] The feature extraction f of the model is a function mapping process, f:LxI→R c *D, the process converts the input image L and the position region I into a vector with a size of cxD through a mapping relationship, and then performs an outer product operation on the outputs of the feature extraction processes f A and f B , so as to obtain the bilinear feature vector at the corresponding position. The pooling function P integrates the obtained bilinear features, and the purpose is to obtain a feature function for fine-grained classification.

[0088] In the pooling process, the bilinear features at each corresponding position on the image are accumulated, and the specific calculation is as follows:

[0089]

[0090] In the formula, represents the bilinear feature accumulation result, l represents a sub-region with an area of the pooling kernel size in the region, and B i represents a bilinear operation.

[0091] According to the relevant guide manual, a standard action library is established for a typical work scene, and key actions are segmented, effective actions are identified, and redundant transition actions are removed to form a key action video sample library. First, the live-line work site video is marked as two video segments, i.e., effective key operation actions and interval transition actions, and the transition perception context network is trained to realize automatic segmentation of the post training monitoring video.

[0092] The transition perception context network mainly includes two core parts, i.e., a timing context detector and a transition perception classifier. The timing context detector takes a standard SSD (single-stage multi-frame detector) framework as a main part, and can encode long-time scale context information by embedding multi-scale bidirectional Conv-LSTM (a deep learning module combining convolutional neural network and long short-term memory network) units in different feature maps. The transition perception classifier classifies actions and action states by tracking and identifying dynamic parts, thereby distinguishing segments in the monitoring video, extracting effective action videos of the work personnel in the monitoring video, and ensuring that the video segments containing complete operations can be detected during action detection.

[0093] Long-term timing context information is crucial for spatio-temporal action detection. However, the standard SSD performs action detection from multiple feature maps of different sizes, and does not consider timing context information. To extract timing context, a Bi-ConvLST unit can be embedded in the SSD to design a cyclic detector for detecting actions. As a kind of LSTM, ConvLSTM can encode long-term information and is more suitable for processing video data, so the ConvLSTM unit can replace the multiplication operation of the fully connected operation in the LSTM unit with a convolution operation, thereby maintaining the spatial structure of the frame over time.

[0094] The specific details of the transition perception classifier include that in the monitoring video, effective operation video segments and transition action video segments often appear alternately, and if the effective operation video segments cannot be accurately extracted from the monitoring video, the task completion of the work personnel cannot be analyzed. The instances in the transition state have similarities with the target actions, and the detection is prone to confusion. The transition perception classifier classifies the actions and action states in the video frame by designing two classifiers, can locate the start frame and end frame of the effective operation video segment in the monitoring video, and thereby divides the entire monitoring video into key operation video segments and transition video segments.

[0095] The multi-scale feature pyramid network is combined with the adaptive attention mechanism to identify abnormal behaviors of the power plant workers in the transition video segments.

[0096] The multi-scale feature pyramid network is combined with an adaptive attention mechanism to identify abnormal operation behaviors of the power plant operators in the transition video segment, including:

[0097] A transition image sequence is extracted from the transition video segment, and a multi-scale feature pyramid network is constructed based on the transition image sequence;

[0098] A sparse optical flow field of the transition image in the topmost layer of the multi-scale feature pyramid network is calculated, and an initial optical flow vector of adjacent transition images is obtained through inter-frame difference processing;

[0099] The optical flow vectors of the remaining transition images are calculated based on the initial optical flow vector, and are fused with the initial optical flow vector to obtain a fused optical flow vector;

[0100] The fused optical flow vector is compared with a preset threshold value, if the fused optical flow vector is greater than the preset threshold value, the optical flow vector of the largest transition image is selected as an abnormal point and is modified to the mean value of the optical flow vector, otherwise, the fused optical flow vector is passed to the next layer until it is passed to the base layer of the multi-scale feature pyramid network;

[0101] The fused optical flow vector of each layer is converted into a depth map, the depth map is spliced based on an attention mechanism, and the abnormal operation behavior is segmented from the spliced channel by using a transition perception context network.

[0102] The posture abnormal behavior of the power plant operator in the behavior video segment is identified by using a time series model.

[0103] S2, acquire the binocular image of the power equipment, construct a multi-stage binocular stereo block matching model, match the feature points in the binocular image of the power equipment, and calculate the safety distance between the power plant operators and the live body based on the matching result.

[0104] The binocular image of the power equipment is acquired, a multi-stage binocular stereo block matching model is constructed, the feature points in the binocular image of the power equipment are matched, and the safety distance between the power plant operators and the live body is calculated based on the matching result, including:

[0105] The binocular image of the power equipment is acquired by using a binocular camera, and a multi-stage binocular stereo block matching model is constructed to match the binocular image, and a disparity image is obtained;

[0106] The depth value of the power equipment point in the disparity image is obtained by using the triangulation method, and the three-dimensional coordinate points of the disparity image are reconstructed by improving the semi-global block matching algorithm;

[0107] Based on the posture abnormal behavior of the power plant operator, the spatial distance between the reconstructed three-dimensional coordinate points and the feature points of the target object is calculated, and the safety distance between the power plant operator and the live body is determined according to the spatial distance.

[0108] wherein, based on the abnormal behavior of the power plant worker, the spatial distance between the reconstructed three-dimensional coordinate point and the target object feature point is calculated, and the safety distance between the power plant worker and the live body is determined according to the spatial distance, including:

[0109] The abnormal behavior of the power plant worker is decomposed into atomic action units, a picture area of a preset size is intercepted with the moving target as the center point, the picture area is divided into small areas of a preset size, and a plurality of prior anchor boxes of different sizes are set at the center points of each small area.

[0110] A power plant worker recognition model is established, and the prior anchor boxes of different sizes are distributed to target regions of different scales according to their sizes through the power plant worker recognition model.

[0111] The target class confidence and the target position confidence of each target region are calculated, and the confidence calculation result is taken as the prediction result of the prior anchor box, and the confidence is corrected according to the existence of the target and the target class in the prior anchor box.

[0112] wherein, the confidence is corrected according to the existence of the target and the target class in the prior anchor box includes:

[0113] A double causal path is established based on the existence of the target, the target class and the pre-defined environmental interference factor in the prior anchor box, and an intervention variable is introduced to test the double causal path.

[0114] A main branch encoder and an auxiliary branch encoder are respectively constructed to calculate the mutual information of the double causal path, and based on the mutual information, a causal intervention is performed in the prior anchor box through the double causal path, and the influence of the environmental interference factor on the confidence is evaluated.

[0115] Based on the influence degree of the environmental interference factor on the confidence, the confidence is corrected by using the introduced interference compensation and the class prior probability, and the corrected confidence is injected back into the main branch encoder, so that the output of the main branch encoder is consistent with the semantics of the corrected confidence.

[0116] Based on the prediction result of the corrected prior anchor box, a non-maximum suppression algorithm is used to remove repeated prediction results as the output result of the power plant worker recognition model.

[0117] Based on the output result of the power plant worker recognition model, the target object feature points are extracted, the three-dimensional coordinates of the target object feature points and the surface centers of the two points of the live body are obtained, and the safety distance between the power plant worker and the live body is calculated.

[0118] According to the safety distance, an electronic fence is drawn in the live body area, and it is monitored whether the power plant worker is in the electronic fence area, if the power plant worker exceeds the electronic fence area, the power plant worker is prompted by voice to re-enter the electronic fence area.

[0119] It should be noted that for the power plant near electricity operation scene, a kind of intelligent safety distance monitoring edge device is developed. Using high-precision binocular camera and intelligent target recognition technology, binocular image stereo matching and other advanced algorithms, image processing, image analysis, target recognition, safety distance calculation are completed on site, and the results are visualized.

[0120] Binocular stereo matching is an important technology in computer vision, which can obtain three-dimensional information of objects through images captured by two cameras. However, traditional binocular stereo matching methods have some problems, such as sensitivity to noise, illumination changes, and low matching efficiency. Therefore, the binocular stereo matching method based on deep learning can improve the matching efficiency and accuracy, and improve the accuracy of image processing.

[0121] The binocular stereo matching method based on deep learning includes the following steps: a multi-stage binocular stereo region matching model based on channel attention module is used for feature extraction of images; a traditional stereo matching algorithm is used to match the features of two images, and a classifier is trained using a live working scene element efficient recognition model to filter and optimize the matching results.

[0122] The power equipment space distance measurement technology based on binocular stereo matching proposed in the application can calculate the spatial distance of the equipment according to the depth map by triangulation method. The depth values of all points in the parallax image are obtained using this technology, and the three-dimensional coordinate points of the image are reconstructed.

[0123] On the basis of the reconstructed three-dimensional coordinate point data, the spatial distance between feature points can be calculated by combining intelligent target recognition technology, further solving the problem of difficult accurate measurement of safety distance for large vehicle overhead operation in substation near electricity operation. By identifying the feature points of vehicles, near electricity operation personnel and live equipment, and calculating the distance between them, accurate understanding and identification of near electricity operation scene can be realized, and operation safety and efficiency can be improved, which specifically includes:

[0124] Binocular images of power equipment are obtained using binocular cameras, then binocular stereo matching is performed to obtain parallax images; the depth values of all points in the parallax image are obtained using triangulation method, and the three-dimensional coordinate points of the image are reconstructed to obtain the spatial information of the power equipment; the feature points of the target objects that need to be measured for spatial distance are extracted by combining intelligent target recognition technology; the spatial distance between the feature points is calculated according to the reconstructed three-dimensional coordinate point data.

[0125] The application takes edge intelligent devices as a detection platform, first perceives substation operation images by using binocular cameras, performs distance detection by using binocular stereo vision technology based on SGBM (Semi-Global Block Matching) stereo matching technology, acquires a dense depth map and a three-dimensional point cloud model, secondly performs operation object detection by using a light-weight substation operator equipment recognition model, and finally performs dynamic 3D distance recognition of substation operators based on a substation operation site dense depth map and target detection results, so as to realize accurate determination and out-of-bound reminding of various distances of operators from live equipment and operation objects.

[0126] The binocular stereo vision matching principle can be simplified as shown in Figure 4 , wherein the imaging planes of left and right cameras are located in the same plane, O L and O R are the optical centers of left and right cameras respectively, two cameras simultaneously observe a same target point P in space, and left and right imaging points P L and P R are obtained respectively, T is the baseline distance of binocular cameras, Z is the distance of the target point from the optical center, generally referred to as depth, and f is the focal length of the camera.

[0127] According to the similar triangle principle, when the coordinates of the imaging points of the target point in left and right images are known, the three-dimensional coordinates of the target point in the actual world coordinates can be obtained, that is:

[0128]

[0129] In the formula, X L and X R respectively represent the coordinates of the same object or the same feature in the left image and the right image respectively; X L -X R represents the parallax; and f represents the focal length of the camera.

[0130] Binocular camera calibration: during the photographing process of the camera, the coordinates of the target point in the world coordinates are converted into two-dimensional coordinates in the image coordinate system while retaining the color information. Conversely, in the three-dimensional reconstruction process of binocular vision, it is necessary to restore the two-dimensional coordinates in the image coordinate system into the three-dimensional coordinates in the world coordinate system. That is, if the positions of the left and right imaging points in the image coordinate system are known, and if it is desired to convert them into three-dimensional information in the world coordinate system, a spatial transformation matrix is required, and the internal and external parameters of the binocular camera are required, and the parameters can be obtained by binocular camera calibration, and the expression is:

[0131]

[0132] , wherein [αβ] is the focal length of the camera; γ is the distortion coefficient of the camera; [u0 vo ] is the principal point coordinate; K is the internal parameter of the camera; R is the rotation matrix of the camera; t is the translation vector of the camera; [R t] is the external parameter position parameter of the camera; (X w , Y w , Z w ) is the three-dimensional point coordinate in the world coordinate system.

[0133] The internal parameter, the external parameter and the distortion coefficient obtained by camera calibration are used for correcting the power operation image, and lay a foundation for subsequent acquisition of three-dimensional point cloud data of live-line operation.

[0134] The SGBM algorithm in the semi-global stereo matching algorithm is widely used because of its high accuracy. The core of the algorithm is as follows: the advantages of the global stereo matching algorithm are used to find the minimum value of the global energy function, so as to obtain the best right matching point of the left imaging point, and the best disparity map is obtained. This method mainly uses mutual information for stereo matching. The algorithm process is divided into the following steps: preprocessing, cost calculation, dynamic programming and post-processing.

[0135] 1) Preprocessing: first, the initial image obtained is processed according to the Sobel operator, and then the image obtained by preprocessing is traversed for each pixel point, and a new image is obtained according to the function. The purpose of preprocessing is to collect the changes of the image in a certain direction, and to prepare for cost calculation. The expression of the function is as follows:

[0136]

[0137] In the formula, pre represents a preset threshold parameter for controlling the mapping behavior of input P; P NEW is the new image obtained by function processing.

[0138] 2) Cost calculation: this part is divided into two parts. The image processed by the Sobel operator is calculated by gradient, and the image not processed is sampled to obtain the SAD cost, and the expression is as follows:

[0139]

[0140] In the formula, SAD cost is the sum of absolute differences; p is the pixel value; (x, y) is the matching pixel coordinate in the left image; d is the candidate disparity; I L , I R are the gray values of the left and right images respectively; k is the radius of the matching window.

[0141] 3) Dynamic programming: the cost aggregation method of SGBM adopts the same global energy optimization strategy as the global matching algorithm. First, the energy function is constructed as follows:

[0142]

[0143] In the formula, C is the total sum of disparity map matching cost, P1 and P2 are two penalty coefficients, respectively, for different disparity difference penalty, and the final purpose is to make the obtained disparity map as smooth as possible; D is a disparity map, p is a current pixel point, N p Indicates the neighborhood of pixel p, T[] is an indicator function (1 when the condition is true, otherwise 0), wherein:

[0144] When the neighborhood pixel q and the disparity of p are different by 1, a fixed penalty P1 is applied, and the main function is to allow the disparity to change slightly (such as a tilted surface).

[0145] When the disparity changes more than 1, a larger penalty P2 (usually P2>>P1) is applied, and the main function is to suppress discontinuous disparity jumps (such as occlusion boundaries).

[0146] 4) Post-processing: After the above dynamic programming, an initial disparity map is obtained, but the disparity map often has large errors and needs to be further processed to optimize the disparity map and improve the depth accuracy of the target point. Generally, the left-right consistency method is used to remove the incorrect matching of the initial disparity map, and according to the fact that each pixel point can only have one correct disparity, the incorrect matching is filtered out.

[0147] However, the SGBM algorithm provided by OPENCV is based on mutual information, although the matching accuracy is high, but the operation amount is often very large, and the hardware requirement for processing image is high, and the final image processing efficiency is not high. In practical application, it is difficult to be widely applied due to poor real-time performance, and Census transformation is widely used in stereo matching field because of its good detection of local features of image and high speed. As shown in Figure 5 Based on the traditional SGBM algorithm, the Census transformation is applied to the initial cost calculation of SGBM, which specifically includes:

[0148] A window with a size of n x m is constructed, and the pixel gray values in the window are compared based on the preset condition with the gray value of the center pixel of the window, and the values obtained by comparison are composed into a bit string, through which the closeness between the matching point pairs can be represented, and the greater the probability of two matching points being the same point, the greater the similarity of the bit string, wherein the expression of the preset condition is:

[0149]

[0150] After obtaining the Census transform value corresponding to the pixel, the matching cost of the matching point cannot be directly displayed, that is, the best matching point cannot be directly selected, and the transform value needs to be further processed. Therefore, based on the obtained transform value, the Hamming distance of the Census transform value of two pixels in the left and right image pairs is calculated, that is, the corresponding bits of the two bit strings are traversed to find the number of differences, which directly reflects the similarity of the two points, that is, the smaller the value, the greater the probability that the two points of the left and right images are the same point in the three-dimensional space. After this step, the matching cost of the entire image can be obtained, and the calculation formula of the Hamming distance is:

[0151]

[0152] In the formula, Q is the Hamming distance calculation result of the two windows; I l , I r are the Census transform bit string values of the left and right images respectively; n is the number of window pixels; ^, >> and & represent exclusive or operation, shift operation and bitwise and operation respectively, Q m is the Hamming distance calculation result of the mth window, and the matching cost of the left and right pixel points is obtained by Hamming distance calculation.

[0153] In order to improve the matching accuracy of the pixel points, the method of quadratic curve interpolation is used to obtain the best pixel accuracy of the target point, that is, through three lowest matching points, the lowest point of the parabola is calculated according to the solution of the binary equation, that is, the extreme point of the curve, and the disparity value corresponding to the point is the new sub-pixel disparity value of the target point, as shown in Figure 6 Through the solution of the sub-disparity value, the accuracy of the depth of the pixel points is greatly improved, wherein the expression of the lowest point of the parabola is calculated according to the solution of the binary equation through three lowest matching points:

[0154]

[0155] The safety distance calculation of live working personnel and equipment: the algorithm processing steps are as shown in Figure 7 The binocular camera is used as the data acquisition device, the improved SGBM stereo matching algorithm is used for real-time three-dimensional reconstruction of the substation site operation, the left view is used as the input of the lightweight target detection algorithm for operation personnel and equipment detection, the corresponding three-dimensional coordinates are obtained for safety distance discrimination, and it is ensured that the power plant personnel and construction equipment are in the operation safety area.

[0156] Considering that the occurrence of live-line work site safety risks has strong correlation with the subjective behavior of the operating personnel, and the initiative behavior of the operating personnel has strong dynamic continuity and instantaneity, and thus the safety risks and the violation behaviors in the electric power industry also occur instantaneously and show dynamicity and strong instantaneity, the application adopts a detection frame of the operating personnel and equipment through target detection, calculates the mean value of the corresponding point cloud data in the detection frame as the three-dimensional surface center of the operating personnel and equipment, and then accurately calculates the spatial distance between the operating personnel and the equipment, and the distance calculation method between the coordinate points comprises:

[0157] Through the calculated three-dimensional coordinates (X k ,Y k ,Z k ) and (X,Y,Z) of the two-point surface centers of the operating personnel and the equipment in the live-line work image, the Euclidean distance between the two points is directly obtained:

[0158]

[0159] There are many moving targets in the live-line work of the power plant, and environmental factors such as equipment action and indicator light state change may be detected as moving targets in addition to the work object, in order to solve this problem, the application establishes a work object recognition model based on the YOLO algorithm, and further confirms the work object information on the basis of the target detection in the previous step.

[0160] After reading the video frame, it is necessary to adjust the size first, and on the basis of the moving target detection result in the foregoing, a picture area of 608x608 size is intercepted as the input with the moving target as the center point, the whole picture is evenly divided into 19x19 small areas, and 9 prior anchor frames of different sizes are set at the center points of each area. Subsequently, the model extracts the picture features through the feature extraction network, and predicts the targets in each area through the prediction head. For each area, since the type of the target to be predicted is only the work object, the number of categories C=1, and for each target, 4 coordinate values and 1 confidence value are needed to be predicted, a total of 5 values, so the prediction result is represented by a 19x19x(9x(5+1)) matrix. Finally, the model adopts non-maximum suppression to screen and remove multiple results of predicting the same target, and finally obtains the work object detection result and position information in the monitoring picture.

[0161] In order to realize multi-scale operation object detection, three different scale feature maps are extracted, and 9 different size prior anchor frames are respectively distributed to different scale target prediction according to size, the three scale feature maps correspond to 19*19, 38*38, 76*76 detection regions respectively, and the prior anchor frame corresponding to each scale is 3, so that the output of each scale is 19*19*18, and each prior anchor frame is only responsible for predicting the target whose center point is located in the region of the anchor frame.

[0162] Whether there is a target in each region is judged by the final confidence, and the background region anchor frame is not discarded in the prediction process, so when calculating the confidence, the possibility of the existence of the target needs to be included, including target class confidence and target position confidence two parts, and the calculation formula is as follows:

[0163]

[0164] In the formula, confidence is the confidence, for each region, if there is a target, the Pr(Object) value of all anchor frames in the region is close to 1, otherwise it is 0, In order to evaluate the intersection over union between the predicted target and the actual target, the larger the intersection over union is, the more accurate the position prediction is. Generally, the prediction result of each anchor frame needs to use the position coordinates (x, y, w, h), the confidence and the softmax vector representing the category information to represent, and the whole network structure is as shown in Figure 8 Figure 8 In the formula, Type is the type of layer, Filters is the number of convolution kernels, Size is the size of convolution kernel, Output is the spatial size of the output feature map of the layer, Convolutional is the convolution layer, Residual is the residual connection, Avgpool is the average pooling layer, Connected is the full connection layer, Softmax is the activation function, which is used for classification output, Global is the global average pooling, YOLO Detection is the YOLO target detection head, and Scale1 / 2 / 3 is three scales.

[0165] When predicting, the existence of the target in the anchor frame and the category of the target need to be considered, so the confidence is further modified as:

[0166]

[0167] In the formula, Pr(Object) is the probability that there is an object in the current position / region; Pr(Class i ) is the probability that the sample belongs to category i, that is, the confidence that the model predicts that the target is the i-th category; Pr(Class i ​Pr(Object) is the possibility of predicting the target data as the i-th class, and since the model only predicts the work object category, Pr(Object) = 1, Pr(Object) and i Pr(Object) = 1, Pr(Object) and respectively represent the possibility of the presence of the target in the region and the calculated intersection over union. Multiplying the three confidences can obtain the final confidence of the predicted target. For the predicted result, part of the invalid target needs to be filtered out through the confidence, and the non-maximum suppression algorithm is used to remove the repeated prediction results, so as to obtain the accurate target detection result.

[0168] S3, in combination with the rule-breaking behavior of the power plant workers and the safety distance, a lightweight knowledge distillation edge device reasoning model is constructed, and the edge device reasoning model is used to monitor the power plant safety hazards in real time.

[0169] Among them, in combination with the rule-breaking behavior of the power plant workers and the safety distance, a lightweight knowledge distillation edge device reasoning model is constructed, and the edge device reasoning model is used to monitor the power plant safety hazards in real time.

[0170] The rule-breaking behavior of the power plant workers and the safety distance are combined to construct a lightweight knowledge distillation edge device reasoning model based on a convolutional network;

[0171] The edge device reasoning model is compressed through operator optimization technology, tensor optimization technology, and data optimization technology;

[0172] Based on the model compression result, the power plant electrical safety hazards are identified, and the performance of the edge device reasoning model is optimized by combining self-supervised contrast learning and meta-learning to adapt to new types of power plant electrical safety hazards;

[0173] The pre-defined teacher model is deployed in the cloud for training, and knowledge is transferred to the lightweight student model through a dynamic adaptation strategy to optimize the reasoning ability of the edge device reasoning model.

[0174] Among them, the pre-defined teacher model is deployed in the cloud for training, and knowledge is transferred to the lightweight student model through a dynamic adaptation strategy to optimize the reasoning ability of the edge device reasoning model.

[0175] The power plant real-time video data and power equipment binocular images are integrated to obtain a sample set, and the sample set is divided into labeled data and input into the teacher model deployed in the cloud and the local lightweight student model;

[0176] The teacher model and the local lightweight student model respectively output the predicted labels, and the predicted samples with clear labels are introduced into the classification loss function to calculate the label difference between the predicted labels of the student model and the true labels;

[0177] The teacher model is adjusted in weight based on the label difference of the student model, and knowledge is transferred to the student model based on the optimized teacher model through a dynamic adaptation strategy, and the teacher model weight is adjusted reversely;

[0178] The teacher model weight updating and knowledge distillation are iteratively performed until the iteration condition is met, and the knowledge alternation between the teacher model and the student model is realized to optimize the inference ability of the edge device inference model.

[0179] It should be noted that, with the development of power internet of things and energy digital transformation, the application scenarios of power edge intelligent terminals are continuously expanding. Due to the intelligent analysis and parallel computing capability of AI chips, the AI chips become one of the mainstream core devices of power edge terminals, and how to realize the power-related hidden danger rapid identification based on convolutional network on the AI chip with limited computing capability needs to use a model compression method. The present application will compress the model from three aspects of operator optimization, tensor optimization and data optimization.

[0180] In the deep learning framework, a computation graph is a common way to represent a program. Currently, the mainstream framework uses a dynamic computation graph, that is, a large number of multi-dimensional tensors flow in the computation graph as intermediate data. The advantage of this computation graph is strong flexibility, but the data exchange amount between operators is large, which is not suitable for the embedded development environment of the edge with low power consumption and low bandwidth. Therefore, the present application adopts a new computation graph optimization scheme, and uses a node to represent a tensor operation and uses an edge to represent the data flow and data dependency relationship between tensors.

[0181] This computation graph composed of nodes and edges provides a global view for the computing task, and also avoids describing how each computing task is implemented. In a graph, a static memory planning pass can process by pre-allocating all intermediate result tensors. This allocation phase is similar to the register allocation pass in traditional compilers. In theory, a computation graph can be converted into a function formula graph. For example, a constant folding pass can be statically executed in the pre-computation stage of the computation graph to save the running time. On the basis of the static computation graph, the present application also uses two methods of operator fusion and data layout conversion to optimize the computation graph.

[0182] Operator fusion: operator fusion is to merge multiple operations of the original computation graph into one. For the multi-neural network computing card architecture based on domestic chips, the optimization method of fusing multiple operations together can obviously reduce the execution time, because the single Kernel function saves the time consumption of writing back the intermediate results to the global memory. Specifically, the present application realizes the fusion of the following operations: addition, subtraction, square root, summation, maximum value, minimum value, and standard two-dimensional convolution.

[0183] In order to verify the effectiveness of operator fusion, the present application carries out operator fusion on the standard two-dimensional convolution operation on the backbone network ResNet50 and the backbone network MobileNet, and gives the time required for a single convolution operation before and after fusion. After operator fusion, the time consumed by a single convolution is reduced by an average of about 30%, which is relatively significant.

[0184] Data layout conversion: Tensor operation is the basic operator of the computation graph, and the operation involved in the Tensor requires different data layout requirements according to different operators. For the domestic neural network calculation rk1808, 4x4 tensor operation is generally used, so the data needs to be cut into 4x4 blocks for storage to optimize the local memory efficiency. The original tensor matrix size is 4x8, in order to adapt to the 4x4 tensor operation, the original tensor needs to be cut into 2x4x2x2 to align the hardware layout of the neural network calculation card, and improve the calculation efficiency and utilization.

[0185] Tensor optimization based on random forest and particle swarm optimization algorithm: Tensor optimization fully considers the collaborative optimization of operation and data when generating target code. In order to optimize memory performance, deep learning processors often design on-chip cache close to the calculation unit, which is much faster than accessing DDR off-chip. In order to simplify the structure design, the deep learning processor does not design the on-chip cache as a software transparent cache structure, so the software programmer needs to manage these on-chip caches explicitly. Reasonable use of on-chip cache can greatly improve program performance, and from the perspective of optimizing operation alone, it cannot be done. In order to fully utilize the on-chip cache in the application scenario of supporting variable size vector multiplication matrix and matrix multiplication matrix operation, the code generator must consider operation optimization and data optimization. On the other hand, since the deep learning processor supports vectorization operation and loop unrolling according to fixed dimensions, from the perspective of improving memory efficiency, off-chip data often needs to be specially aligned and dimensionally transformed.

[0186] First step, random forest and particle swarm optimization algorithm:

[0187] Random forest is a supervised learning algorithm based on Bagging and decision tree. Decision tree is a widely used classification and regression method, which has high accuracy in most cases. However, when the data is complex or noisy, decision tree is prone to overfitting, which reduces the accuracy. Random forest is an ensemble learning of decision tree, which is not easily disturbed by noise, greatly improves the problem of overfitting of decision tree, achieves good fitting effect on high-dimensional data, and can operate in parallel, with fast training speed and high efficiency.

[0188] Bagging (Bootstrap aggregating) is an ensemble learning method based on the Bootstrap idea. It randomly selects training subsets from the original training sample set with replacement, trains multiple regression trees, and generates a random forest by repeating the process. The final prediction result is determined by the average of all tree prediction values. When generating training subsets using the Bagging method, the probability of each sample in the sample set D not being selected is (1-1 / N)N, where N is the number of samples in the sample set. When N is large enough, (1-1 / N)N converges to 0.368, so in each sampling process, 37% of the data in the sample set will not be used to train the random forest regression model. These samples are called out-of-bag (OOB) data. OOB estimation is used to evaluate the fitting effect of the model.

[0189] The steps to establish a random forest regression model are as follows:

[0190] 1) Use the Bootstrap method to randomly select n (n<N) samples from N original training samples with replacement to generate m training subsets;

[0191] 2) Use the training subset to train the regression tree, randomly select a subset of sample features at the node, and perform regression tree left and right subtree division based on the minimum mean square error. Recursively build the tree until the termination condition is met;

[0192] 3) Repeat the above steps to form a random forest of multiple regression trees;

[0193] 4) Input the test sample into the random forest regression model, take the average of all regression tree prediction values as the final prediction result, and compare it with the actual value to evaluate the fitting effect of the model.

[0194] Particle swarm optimization algorithm is a global search algorithm based on cooperation and competition between individuals. The system is initialized as a set of random solutions, called particles. Optimization is achieved by the flight of particles in the search space, which is iteration in mathematical formula. It does not have crossover and mutation operators like genetic algorithms, but particles search for the best particle in the solution space. The basic principle is as follows:

[0195] In a search space with D variables, the position and velocity of a particle can be represented as:

[0196] x i (t)=[x i,1 (t),x i,2 (t),…,x i,D (t)],i=1,2,3,…N;

[0197] v i (t)=[v i,1(t), v i,2 (t), …, v i,D (t)], i = 1, 2, 3, … N;

[0198] where N is the total number of particles, x i,D (t) and v i,D (t) are the position and velocity of the D-dimensional variable in the i-th particle at the t-th iteration, respectively.

[0199] When the iteration optimization reaches the t-th time, the optimal position of the current population and the variable state are pbest i (t) and P i,D (t), respectively. At this time, the global optimal position is represented by gbest i (t), and the optimal state of each variable in the particle is g i,D (t). Specifically, it is represented as:

[0200] pbest i (t) = [P i,1 (t), P i,2 (t), …, P i,D (t)], … i = 1, 2, 3, …, N;

[0201] gbest i (t) = [g i,1 (t), g i,2 (t), …, g i,D (t)], … i = 1, 2, 3, …, N;

[0202] The update of the particle position in the multi-objective particle swarm algorithm is mainly based on the update of the velocity, and the velocity is mainly updated according to pbest i (t) and gbest i (t). The specific update expression of the t+1-th iteration is:

[0203] v i (t+1) = ω(t) · v i (t) + c1 · rand · (pbest i (t) - x i (t)) + c2 · rand() · (gbest i (t) - x i (t));

[0204] x i (t+1) = x i (t) + v i (t+1);

[0205] In the formula, w is an inertial weight value, rand is a random number between 0 and 1, and c1 and c2 are both learning factors.

[0206] The speed expression can be divided into three parts, the first part is the previous speed of the particle, the second part is the "cognition" part, which represents the thinking of the particle itself, and the third part is the "social" part, which represents the information sharing and mutual cooperation between particles.

[0207] The second step, the main calculation steps of the PSO algorithm are as follows:

[0208] 1) Initialize the particle group, set the acceleration constants c1 and c2, the maximum evolution generation, and set the current evolution generation to 1, and randomly generate the positions and speeds of m particles in the defined space;

[0209] Evaluate the population, calculate the fitness value of each particle in each dimension of the space;

[0210] 2) Compare the fitness value of the particle with the optimal value pbest of itself. If the current value is better than pbest, set pbest as the current value, that is, update the historical best position with the current position;

[0211] 3) Compare the fitness value of the particle with the optimal value gbest of the population. If the current value is better than gbest, update the global best position with the current position;

[0212] Update the historical position and speed according to the formula to generate a new population;

[0213] 4) Check the end condition. If it is satisfied, end the optimization; otherwise, go to step 2. The end condition is usually that the optimization reaches the maximum evolution generation or the evaluation value is less than a given precision.

[0214] Objective function:

[0215] In order to realize automatic code generation, the present application uses tensor expressions to specify tensor operators. Let ε be the tensor expression space. The tensor expression retains many underlying implementation details, such as loop order and memory range, etc.

[0216] The third step, the cost generation module based on random forest:

[0217] The random forest prediction method based on machine learning theory is used to predict the hardware runtime cost of the low-level code, and the random forest prediction method extracts specific domain features from the given low-level abstract syntax tree. These features include loop structure information (such as memory access count and data reuse rate) and general annotations (such as vectorization, unrolling, and thread binding).

[0218] For a given data set D={(e i ,s i ,ci The statistical cost model is trained while considering that only the relative order of program running time is concerned in the selection process, rather than their absolute values, so the expression of the training target function is:

[0219]

[0220] wherein c1 and c2 are acceleration constants respectively, is the hardware running time cost, x i , x j are both two instance samples used to compare their relative scores.

[0221] The fourth step is an optimization module based on a particle swarm optimization algorithm:

[0222] Since the optimization space is too large, the optimization module of the present application uses a particle swarm optimization algorithm to control the search loop with the hardware running time cost predicted by the cost generation module as a target function, and the specific process is shown in the algorithm pseudo code Table 1. In each iteration, it must select a batch of candidate programs according to the hardware running time cost predicted by the cost generation module, and query the actual hardware running time cost f(x) on the actual hardware.

[0223] Table 1: Optimization module based on particle swarm optimization algorithm

[0224]

[0225] Specifically, the present application uses batch parallel Markov chains to improve the prediction throughput of the statistical cost model. The present application selects a batch of candidates with the best performance to run on the real hardware, and the performance data collected is used to update The present application uses the state of the Markov chain to remain unchanged between updates. Random selection is used with greedy selection (b=0.05) to ensure exploration. At the same time, the present application uses diversity-aware exploration, and considers quality and diversity when selecting b candidates for evaluation. Assuming that the planned configuration s can be decomposed into m components s=[s1, s2, …s m ], the expression for maximizing the target to select a candidate set S from the candidate set λb is:

[0226]

[0227] wherein the first term gives priority to candidates with low running time cost, the second term calculates the number of different configuration components covered by S, L(S) is a sub-module function, and a greedy algorithm is used to obtain an approximate solution, s is the planned configuration, m is the number of configuration components, and s jFor the jth component, a is a regularization coefficient balancing the importance of the two parts, and g(e,s) is a feature constructor.

[0228] Data optimization:

[0229] Data optimization is the optimal data layout required by the AI chip hardware platform adopted according to the present application, which performs secondary optimization of model data for the hardware platform. The main optimization includes:

[0230] 1) Align the data according to the requirements of the operation to speed up the memory access. The vector operation of the deep learning processor has certain requirements for data alignment, and aligned access can speed up memory access.

[0231] 2) Perform dimension transformation on the model data according to the requirements of the operation to improve the memory locality. Convolution algorithms require data to be interpreted as multi-dimensional arrays, and the arrangement order of multi-dimensional arrays will affect the number of memory jumps. In order to minimize memory jumps, data needs to be dimensionally transformed.

[0232] 3) Precision tuning to improve operation speed. Deep learning processors often support low-precision operations, such as 16-bit quantization and 8-bit quantization, and different precision operations have different performance. Therefore, during compilation, model data can be optimized for precision to improve performance.

[0233] The working principle of the data optimizer is shown in Figure 9 , which first performs the above-mentioned three optimizations on the model data according to the requirements of the data layout of the hardware platform code; packs the hardware platform code and the optimized model data to generate a deployable model.

[0234] The main purpose is to achieve precision tuning and improve running speed through neural network model quantization. First, the model is quantized according to the chip situation, and then the deep neural network model compression technology is used to compress the model volume as much as possible under the premise of minimizing precision loss, greatly reducing the model parameter quantity and calculation quantity, and speeding up the model inference speed, which is helpful for model deployment in power edge intelligent devices with limited computing resources, plays a role in real-time monitoring of power operation site, and can effectively improve the monitoring quality of on-site operation personnel and reduce the safety risk of on-site operation personnel.

[0235] Due to the performance limitations of the computing hardware devices adopted by the present application, the size of the network model and the calculation requirements must also be optimized accordingly to adapt to device use. The present application first quantizes the model.

[0236] The implementation of neural network quantization requires converting common operations (convolution, matrix multiplication, activation function, pooling, splicing, etc.) into equivalent 8-bit integer versions of operations, and then adding quantization and dequantization operations before and after the operations, respectively. The quantization operation is to convert the input from a floating-point number (float) to an 8-bit integer (uint8), and the dequantization operation is to convert the output from an 8-bit integer back to a floating-point number. Quantization is actually an operation to reduce the number of value bits, which is a way to represent a numerical value. Currently, the weights of neural network training are stored in 32-bit floating-point numbers, and neural networks have millions or even more weights connected between neurons, so they will occupy a large amount of storage space. If these models are run on a PC, there will be no obstacles due to the powerful computing performance and storage resources of the PC. However, if they are placed on an embedded end, it is not realistic because the computing performance and storage resources of the PC are often tens of times that of the embedded end. However, if it is simply rounded off or converted to an integer in one step, it will have a devastating impact on model accuracy. It is expected that a large number of parameter weights and feature maps will be represented using 8-bit integer data format without much loss of model accuracy. Even if a lower bit width such as 4 bits, 2 bits, or 1 bit is used, similar performance to the floating-point number model can be achieved.

[0237] The most significant advantage of the quantization method is that it can greatly reduce the bandwidth and storage space, as shown in the table. For example, using 8-bit integer parameter weights and feature maps can reduce the bandwidth requirement by 4 times compared to a 32-bit floating-point number model when performing calculations. In addition, compared to floating-point numbers, integer numbers are faster, have lower power consumption, and require less memory storage when performing operations.

[0238] Table 2: Comparison of 8-bit integer and 32-bit floating-point operations in terms of power consumption and storage space

[0239]

[0240] Taking the Relu activation function as an example, the Relu operation before quantization is as shown in Figure 10 (in the Figure 10 , Input is the input, Relu is the Relu activation function, and Output is the output), and the Relu operation after network quantization is as shown in Figure 11As shown, Input (float) is a neural network feature value in float format and serves as input; Min / Max is the minimum value / maximum value for calculating the dynamic range of input data, Quantize is used to convert float data (float) into integer representation, Eight Bit is an 8-bit integer data, indicating that the quantized value is an 8-bit integer, QuantizedRelu is a quantized ReLU activation function, which applies ReLU activation in the integer space: all negative numbers are set to zero, and non-negative numbers are retained. Unlike ordinary ReLU, it operates within the quantized value range, Dequantize is used to map the 8-bit integer value back to the float value, and restore it to the approximate original feature, Output (float) is the output. For a convolutional neural network model, there are two types of parameters that need to be quantized, one is the weight of the model, and the other is the activation value of the model, that is, the feature map. As shown in Figure 12 If the weight range of a certain layer is [a, b], the quantization method is: the minimum value of the weight is taken as a, and the maximum value of the weight is taken as b. The weight is fine-tuned through training to map the weight range to [0, 255], so that it becomes an 8-bit integer. As for the activation value, its value is largely dependent on the input of the network. Assuming that the activation value range is still [a, b] during the retraining phase, the activation value is updated using the EMA (Exponential moving average) with a smoothing parameter close to 1 during training. The role of EMA is to slow down the trend when the activation value range changes dramatically.

[0241] According to another embodiment of the present application, as Figure 2 As shown, a lightweight image recognition system for power plant safety is also provided, which comprises:

[0242] A violation behavior recognition module 1 is configured to acquire real-time video data of the power plant, construct a multi-scale adaptive perception image anomaly detection model, and identify the violation behavior of the power plant operating personnel in the real-time video data of the power plant by using the image anomaly detection model.

[0243] A safety distance calculation module 2 is configured to acquire binocular images of power equipment, construct a multi-stage binocular stereo region matching model to match feature points in the binocular images of the power equipment, and calculate the safety distance between the power plant operating personnel and the live body based on the matching result.

[0244] A safety hazard monitoring module 3 is configured to combine the violation behavior of the power plant operating personnel and the safety distance, construct a lightweight knowledge distillation edge device reasoning model, and monitor the safety hazards of the power plant in real time by using the edge device reasoning model.

[0245] The above merely provides the preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A lightweight image recognition method for power plant safety, characterized by, The method comprises: acquiring power plant real-time video data, constructing a multi-scale adaptive perception image anomaly detection model, and identifying the rule violation behavior of the power plant operating personnel in the power plant real-time video data by using the image anomaly detection model; acquiring power equipment binocular images, constructing a multi-stage binocular stereo area matching model to match feature points in the power equipment binocular images, and calculating the safety distance between the power plant operating personnel and the live body based on the matching result; combining the rule violation behavior of the power plant operating personnel and the safety distance, constructing a lightweight knowledge distillation edge device reasoning model, and monitoring the power plant safety hazards in real time by using the edge device reasoning model; The method comprises: using a bilinear convolutional neural network to mine visual features and spatial topological features in the power plant real-time video data, and fusing the visual features and spatial topological features to identify a reference power operation scene; based on the reference power operation scene, using a time context detector in a transition perception context network to encode context information, and using a transition perception classifier to segment the power plant real-time video data to form behavior video segments and transition video segments; combining a multi-scale feature pyramid network and an adaptive attention mechanism to identify abnormal operation behaviors of the power plant operating personnel in the transition video segments; using a time series model to identify posture abnormal behaviors of the power plant operating personnel in the behavior video segments; The method comprises: acquiring binocular images of power equipment by using a binocular camera, constructing a multi-stage binocular stereo area matching model to match the binocular images, and obtaining a parallax image; using a triangulation method to obtain the depth value of the power equipment points in the parallax image, and reconstructing the three-dimensional coordinate points of the parallax image by using an improved semi-global block matching algorithm; based on the posture abnormal behaviors of the power plant operating personnel, calculating the spatial distance between the reconstructed three-dimensional coordinate points and the feature points of the target object, and determining the safety distance between the power plant operating personnel and the live body according to the spatial distance; determining an electronic fence in the live body area according to the safety distance, and monitoring whether the power plant operating personnel is in the electronic fence area, if the power plant operating personnel exceeds the electronic fence area, then prompting the power plant operating personnel to re-enter the electronic fence area through voice; The method comprises: constructing a lightweight knowledge distillation edge device reasoning model based on a convolutional network according to the rule violation behavior of the power plant operating personnel and the safety distance; compressing the edge device reasoning model by using operator optimization technology, tensor optimization technology, and data optimization technology respectively; based on the model compression result, identifying power-related hazards in the power plant, and optimizing the performance of the edge device reasoning model by using self-supervised contrast learning and meta-learning to adapt to new types of power-related hazards in the power plant. The pre-defined teacher model is deployed in the cloud for training, and knowledge is transferred to the lightweight student model through a dynamic adaptation strategy to optimize the inference ability of the edge device inference model.

2. The lightweight image recognition method for power plant safety according to claim 1, characterized in that, The combination of the multi-scale feature pyramid network and the adaptive attention mechanism to identify abnormal behaviors of power plant operators in transition video clips includes: extracting a transition image sequence from the transition video clip, and constructing a multi-scale feature pyramid network based on the transition image sequence; calculating the sparse optical flow field of the transition image in the top layer of the multi-scale feature pyramid network, and performing inter-frame difference processing on adjacent transition images to obtain initial optical flow vectors of the adjacent transition images; based on the initial transport optical flow vector, calculating the optical flow vectors of the remaining transition images, and fusing them with the initial optical flow vector to obtain a fused optical flow vector; comparing the fused optical flow vector with a predetermined threshold, if the fused optical flow vector is greater than the predetermined threshold, then iteratively selecting the optical flow vector of the largest transition image as an abnormal point and modifying it to the mean value of the optical flow vector, otherwise, the fused optical flow vector is passed to the next layer until it is passed to the bottom layer of the multi-scale feature pyramid network; convert the fused optical flow vector of each layer into a depth map, perform channel splicing on the depth map based on the attention mechanism, and use the transition perception context network to segment the abnormal behavior from the spliced channels. 3.The power plant security-oriented lightweight image recognition method of claim 1, wherein Based on the abnormal behavior of the power plant operator, the spatial distance between the reconstructed three-dimensional coordinate points and the target object feature points is calculated, and the safe distance between the power plant operator and the live body is determined according to the spatial distance. The abnormal behavior of the power plant operator is decomposed into atomic action units, and a picture area of a predetermined size is intercepted with the moving target as the center point. The picture area is divided into small areas of a predetermined size, and a number of prior anchor boxes of different sizes are set at the center points of each small area. An operator identification model is established, and different sizes of prior anchor boxes are distributed to different target regions according to their size through the operator identification model. Calculate the target class confidence and target position confidence of each target region, and use the confidence calculation result as the prediction result of the prior anchor box. The confidence is corrected according to the existence of the target and the target class inside the prior anchor box. Based on the prediction result of the corrected prior anchor box, use the non-maximum suppression algorithm to remove duplicate prediction results as the output result of the operator identification model. Based on the output result of the operator identification model, extract the target object feature points, and obtain the three-dimensional coordinates of the two points on the surface center of the target object and the live body, and calculate the safe distance between the power plant operator and the live body.

4. The lightweight image recognition method for power plant safety according to claim 3, characterized in that, The confidence is corrected according to the existence of the target and the target class inside the prior anchor box includes: Based on the existence of the target, the target class and the pre-defined environmental interference factor inside the prior anchor box, a double causal path is established, and an intervention variable is introduced to test the double causal path. Respectively construct the main branch encoder and the auxiliary branch encoder to calculate the mutual information of the double causal path, and based on the mutual information, perform causal intervention inside the prior anchor box through the double causal path to evaluate the influence of environmental interference factors on confidence. Based on the influence degree of environmental interference factors on the confidence, the confidence is corrected by introducing interference compensation and category prior probability, and the corrected confidence is injected into the main branch encoder in reverse, so that the output of the main branch encoder is consistent with the semantics of the corrected confidence.

5. The lightweight image recognition method for power plant security according to claim 1, characterized in that, The pre-defined teacher model is deployed in the cloud for training, and a dynamic adaptation strategy is used to transfer knowledge to the lightweight student model to optimize the inference ability of the edge device inference model. The power plant real-time video data and power equipment binocular image are integrated to obtain a sample set, and the sample set is divided into labeled data and input into the teacher model deployed in the cloud and the local lightweight student model. The teacher model and the local lightweight student model output prediction labels, and the prediction samples with clear labels are introduced into a classification loss function to calculate the label difference between the prediction labels of the student model and the true labels. Based on the label difference of the student model, the weight of the teacher model is adjusted, and based on the optimized teacher model, knowledge is transferred to the student model through a dynamic adaptation strategy, and the teacher model weight is adjusted in reverse. The teacher model weight updating and knowledge distillation are iteratively executed until the iteration condition is met, the knowledge between the teacher model and the student model is alternated, and the inference ability of the edge device inference model is optimized.

6. The lightweight image recognition method for power plant safety according to claim 3, characterized in that, The confidence correction formula is: ; wherein denotes the probability that the sample belongs to class i; denotes the likelihood that the prediction target data is of the i-th class; denotes the likelihood that a target is present within the region; denotes the evaluation of the intersection over union between the predicted target and the actual target.

7. A lightweight image recognition system for power plant security, characterized by, The system for implementing the power plant safety-oriented lightweight image recognition method of any one of claims 1-6 comprises: A violation behavior recognition module is configured to obtain power plant real-time video data, construct a multi-scale adaptive perception image anomaly detection model, and identify the violation behavior of power plant workers in the power plant real-time video data using the image anomaly detection model. A safety distance calculation module is configured to obtain power equipment binocular images, construct a multi-stage binocular stereo area matching model to match feature points in the power equipment binocular images, and calculate the safety distance between the power plant workers and live bodies based on the matching results. A safety hazard monitoring module is configured to combine the violation behavior of the power plant workers and the safety distance, construct a lightweight knowledge distillation edge device inference model, and use the edge device inference model to monitor the power plant safety hazards in real time.

Citation Information

Patent Citations

  • Binocular stereo matching detection method and device based on channel attention mechanism

    CN115965578A

  • Binocular vision-based power transformation operator near-electricity distance detection method

    CN116563386A