Energy station abnormal behavior early warning method based on multi-source image recognition and deep learning

By dividing key monitoring areas in the energy station and utilizing multi-source image recognition and deep learning technologies for image acquisition, processing, and feature extraction, combined with improved algorithms and network models, accurate identification and timely early warning of personnel behavior were achieved, solving key issues in the safety management of the energy station.

CN116758475BActive Publication Date: 2025-12-19HANGZHOU YINGJI POWER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310688572.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-09
Publication Date
2025-12-19
Estimated Expiration
2043-06-09

AI Technical Summary

Technical Problem

How to accurately identify personnel behavior in integrated energy stations, prevent unauthorized personnel from entering key areas to operate equipment, and ensure the safe and reliable operation of equipment is a challenge that current technologies struggle to achieve with high accuracy and timely warnings.

Method used

By employing multi-source image recognition and deep learning methods, key monitoring areas are divided at the energy station, multiple image acquisition devices are set up, and image preprocessing, fusion, and feature extraction are performed. Combined with the improved YOLOv5 algorithm and DeepSort target tracking algorithm, personnel targets are identified and tracked. Generative adversarial networks and BiLSTM networks are used for abnormal behavior recognition and decision fusion to achieve accurate early warning.

Benefits of technology

It improved the accuracy of abnormal behavior identification, provided timely warnings, avoided accidents and economic losses, and ensured the safe and reliable operation of the energy station.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758475B_ABST
    Figure CN116758475B_ABST
Patent Text Reader

Abstract

The application discloses an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning, comprising the following steps: dividing an energy station into multiple key monitoring areas, starting a first image acquisition device to patrol the key monitoring areas through scheduling, collecting personnel entering area images, judging whether personnel enter or not, after judging that personnel enter, starting a second and a third image acquisition device to collect multi-source images; processing the collected multi-source images to obtain multi-source fusion images, and adopting an improved YOLOv5 algorithm to detect personnel targets and adopting a DeepSort algorithm to track personnel targets; after extracting face features, behavior trajectory features, skeleton features and fusing the features, inputting the features into a trained generative adversarial network and a BiLSTM network respectively to identify energy station abnormal behaviors and fuse decisions, obtaining an abnormal behavior category of the key monitoring areas of the energy station, and early warning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of comprehensive energy stations, and particularly relates to an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning. BACKGROUND

[0002] In the construction of a comprehensive energy station in a comprehensive park or development zone, multiple energies (gas, water, steam, heat energy, and electric energy) and multiple energy supply and energy saving equipment (generator sets, waste heat boilers, refrigerating units, solar panels, etc.) are combined, and production and energy transmission are centrally managed and controlled through the Internet to create an independent energy island and realize resource integration and comprehensive utilization of regional energy.

[0003] The Internet of Things is a fusion application of intelligent sensing, identification technology, and ubiquitous computing and ubiquitous network. In the 5G era of the Internet of Everything, the Internet of Things uses different types of sensors to perceive surrounding objects and physical environments to provide a basis for data analysis in the application layer of the Internet of Things and gradually realize unattended operation of an energy station. The types of equipment in an energy station are diverse, including heat source units, heat stations, fresh air units, air conditioning units, water pumps, valves, cooling towers, key pipeline fittings, and key electrical equipment. The safe and reliable operation of these devices plays a crucial role in the entire energy station. Therefore, how to accurately identify the behavior of personnel in an energy station, avoid illegal personnel from entering key areas and performing illegal operations on equipment, and thus cause accidents and economic losses, or accurately determine responsibility after an accident is a problem that needs to be solved urgently.

[0004] Based on the above technical problems, a new energy station abnormal behavior early warning method based on multi-source image recognition and deep learning needs to be designed. SUMMARY

[0005] The present application solves the technical problem of overcoming the shortcomings of the prior art, combining the daily management habits and potential safety hazards of an energy station to provide an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning, which can improve the accuracy of abnormal behavior recognition through multi-source image acquisition fusion, personnel detection and tracking, multi-type feature extraction, and abnormal behavior recognition, and timely warning to avoid accidents and economic losses and ensure the safe and reliable operation of the energy station.

[0006] To solve the above technical problems, the technical solution of the present application is:

[0007] The present application provides an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning, comprising:

[0008] S1, divide the energy station into a plurality of key monitoring areas, and a first image acquisition device, a second image acquisition device and a third image acquisition device are arranged in each key monitoring area respectively;

[0009] S2, start the first image acquisition device to patrol the key monitoring area by scheduling, collect the image of personnel entering the area, and judge whether personnel enter, after judging that personnel enter, start the second image acquisition device and the third image acquisition device to collect multi-source images;

[0010] S3, image preprocessing, image registration, image fusion and fusion evaluation are performed on the collected multi-source images to obtain a multi-source fusion image of the key monitoring area;

[0011] S4, the improved YOLOv5 algorithm is used for key monitoring area personnel target detection on the multi-source fusion image of the key monitoring area, and the DeepSort target tracking algorithm is used for key monitoring area personnel target tracking;

[0012] S5, after personnel image information of key monitoring area personnel target detection and tracking, facial feature extraction, behavior trajectory feature extraction and skeleton feature extraction are performed, and feature fusion is performed;

[0013] S6, after the feature fusion feature is input into the trained generation adversarial network and BiLSTM network, the energy station abnormal behavior recognition and decision fusion are performed, and the abnormal behavior category of the key monitoring area of the energy station is obtained; the abnormal behavior includes personnel abnormal wandering behavior, hesitation and retention behavior, violent destruction behavior and illegal operation behavior;

[0014] S7, according to the abnormal behavior category of the key monitoring area of the energy station, the abnormal behavior level where the abnormal behavior is located is analyzed, and the corresponding degree of behavior warning is performed.

[0015] Further, the S1 divides the energy station into a plurality of key monitoring areas, including: based on the key equipment in the energy station including heat source unit, heat station, fresh air unit, air conditioning unit, water pump, valve, cooling tower, key pipeline pipe, main electrical equipment, the energy station is divided into a plurality of key monitoring areas;

[0016] The first image acquisition device is arranged at the entrance and exit of each key monitoring area; the second image acquisition device and the third image acquisition device are arranged according to the different shooting angles of the positions near the key equipment; the resolution of the second image acquisition device and the third image acquisition device is higher than that of the first image acquisition device; the types of the second image acquisition device and the third image acquisition device are different, and the number is at least one.

[0017] Further, the S2, by scheduling starting the first image acquisition device to patrol the key monitoring area, collecting the image of the personnel entering the area, judging whether there is personnel entering, after judging that there is personnel entering, scheduling starting the second image acquisition device and the third image acquisition device to collect multi-source images, comprising:

[0018] Classifying the first image acquisition device, the second image acquisition device and the third image acquisition device;

[0019] First, the first image acquisition device is started by scheduling to patrol each key monitoring area, collect the image of the personnel entering and exiting the area, judge whether there is personnel entering the key monitoring area, if there is personnel entering, calculate the direction of the personnel advancing according to the video image of the personnel entering the area, and then start the second image acquisition device and the third image acquisition device corresponding to the key monitoring area according to the position of the detected area where the personnel enters, extract the video stream image, and collect multi-source images of personnel behavior.

[0020] Further, in the S3, the image preprocessing includes stretching the gray scale of the image by using gamma correction and suppressing the noise of the image by using the Gaussian curvature filtering algorithm;

[0021] The image registration includes first detecting, describing and matching feature points of the images to be registered by using the scale invariant feature transformation (SIFT) algorithm, then deleting the feature vectors with feature values less than a set value in the described feature vectors by using the singular value decomposition (SVD) algorithm to reduce the dimension of the feature vectors, finally reconstructing the feature point description vectors and performing initial matching of the feature points by using the descriptor reorganization strategy, and eliminating the wrong matching points in the matching process to realize image registration;

[0022] The image fusion includes decomposing the multi-source images into high-frequency and low-frequency images by multi-scale transformation, inputting the high-frequency components and low-frequency components into a neural network model for fusion to obtain high-low frequency fusion images, and obtaining the fusion image after inverse transformation;

[0023] The fusion evaluation includes subjective evaluation and objective evaluation; the subjective evaluation is to directly evaluate the overall fusion image by observing the image texture details, clarity and overall information amount with the human eye; the objective evaluation is based on the information entropy, standard deviation, average gradient, mutual information, structural similarity and edge information retention value indexes of the image.

[0024] Further, the S4, the improved YOLOv5 algorithm is used for key monitoring area personnel target detection on the multi-source fusion image of the key monitoring area, comprising:

[0025] Improving the YOLOv5 algorithm: in the network structure of the original YOLOv5 model, CBAM attention mechanism modules are added at the Backbone backbone network, the Neck unit and the Prediction unit, and Ghost lightweight convolution layers are introduced to replace the general convolution layers in the Backbone backbone network, GhostBottleneck structures are introduced to replace the CSP structure in the Backbone, and weighted bidirectional feature pyramids are introduced to replace the bottom-up feature pyramids in the Neck unit;

[0026] Labeling the multi-source fusion image of the key monitoring area, obtaining the labeled data set, and training the improved YOLOv5 algorithm model to obtain a personnel target detection model;

[0027] The multi-source fusion image to be detected in the key monitoring area is input into the personnel target detection model to detect the personnel and position in the key monitoring area, and the positioning frame of the personnel is output.

[0028] Further, the personnel target tracking in the key monitoring area using the DeepSort target tracking algorithm comprises:

[0029] The Kalman filter algorithm is used to predict the position and state of the personnel target in the next frame;

[0030] The Hungarian algorithm is used for matching to obtain the trajectory of the personnel before and after the video image, and the Mahalanobis distance between the personnel detection frame and the personnel tracking frame is calculated, when the distance is less than a set threshold, the two are associated with each other, and the matching is successful; otherwise, the personnel position is recalculated.

[0031] The Kalman filter update formula is used to update the tracking frame parameters that have been matched, and the personnel target in the next moment is predicted; when the result predicted by the updated parameters cannot be matched, it indicates that the personnel target in the key monitoring area has been lost, and the tracking frame is deleted.

[0032] Further, S5, after the personnel image information after the personnel target detection and tracking in the key monitoring area is extracted, the behavior trajectory feature extraction and the skeleton feature extraction are performed, and the features are fused, comprising:

[0033] The personnel image information after the personnel target detection and tracking in the key monitoring area is input into the pre-trained graph convolutional neural network model for face feature extraction, and the face information in the database is compared, if the comparison is successful, it indicates that the personnel is a normal worker; otherwise, it indicates that the personnel is an abnormal personnel;

[0034] a trajectory model including a person position feature, a direction feature and a speed feature; after the position of each trajectory point is extracted from the person image information, the differential data in the position direction is calculated, and the direction and speed information of the person target are represented by the differential data; the position feature, the direction feature and the speed are input into a pre-trained graph convolutional neural network model for behavior trajectory feature extraction;

[0035] a skeleton sequence is extracted from the person image by using a skeleton extraction algorithm OpenPose, and a skeleton feature is obtained by using a pre-trained graph convolutional neural network model to extract a spatio-temporal skeleton graph G(V, E); V is a person skeleton joint node; E is an edge of the skeleton graph; the skeleton sequence includes N joint nodes in each frame;

[0036] the extracted face feature, behavior trajectory feature and skeleton feature are fused to form a face-trail-skeleton fusion feature.

[0037] Further, the graph convolutional neural network is an attention-enhanced graph convolutional neural network model, which includes one network full connection layer, one long short-term memory network, three attention-enhanced graph convolutional long short-term memory networks and two full connection layers; during the training of the graph convolutional neural network model, the fusion feature is taken as the input data of the model, the cross-entropy loss and attention regularization loss of the network are calculated through each layer of the model, the network parameters are updated by using back propagation, and the training is ended when the training number reaches a set value or the network loss value is lower than a preset threshold.

[0038] Further, in S6, after the fused features are input into the trained generative adversarial network and BiLSTM network respectively for energy station abnormal behavior recognition and decision fusion, the abnormal behavior category of the key monitoring area of the energy station is obtained, including:

[0039] the fused features are input into the trained generative adversarial network for energy station abnormal behavior recognition, and an energy station abnormal behavior recognition result I is obtained;

[0040] the fused features are input into the trained BiLSTM network for energy station abnormal behavior recognition, and an energy station abnormal behavior recognition result II is obtained;

[0041] after the energy station abnormal behavior recognition I and the energy station abnormal behavior recognition result II are decision fused, the abnormal behavior category of the key monitoring area of the energy station is obtained;

[0042] The generative adversarial network includes a generator network, a discriminator network and an optical flow estimation network.

[0043] The BiLSTM network comprises a BiLSTM neural network layer and an output layer, wherein the BiLSTM neural network layer comprises an input layer, a forward propagation layer and a backward propagation layer, the number of nodes of the input layer, the forward propagation layer, the backward propagation layer and the output layer is set in advance, and the weights between nodes of adjacent layers are randomly set.

[0044] Further, according to the abnormal behavior category of the key monitoring area of the energy station, the abnormal behavior level in which the abnormal behavior is located is analyzed, and a corresponding degree of behavior warning is performed, comprising: according to the abnormal behavior category of the key monitoring area of the energy station, the abnormal behavior is matched with the pre-set abnormal behavior level to obtain the abnormal behavior level in which the abnormal behavior is located, if it is a first-level abnormal behavior, sound and light alarm should be performed, and immediately reported to the superior department, the running state and parameters of the monitoring equipment are monitored, and remote control is adopted; if it is a second-level abnormal behavior, sound and light alarm should be performed, and people should be sent to the area to check; if it is a third-level abnormal behavior, sound and light alarm should be performed.

[0045] The beneficial effects of the present application are:

[0046] The present application divides the energy station into multiple key monitoring areas, and respectively sets first image acquisition devices, second image acquisition devices and third image acquisition devices in each key monitoring area; through dispatching to start the first image acquisition device to patrol the key monitoring area, the image of personnel entering the area is collected to judge whether personnel enter, after judging that personnel enter, the second image acquisition device and the third image acquisition device are dispatched to start to collect multi-source images; the collected multi-source images are subjected to image preprocessing, image registration, image fusion and fusion evaluation to obtain multi-source fusion images of the key monitoring area; the improved YOLOv5 algorithm is used for key monitoring area personnel target detection on the multi-source fusion images of the key monitoring area, and the DeepSort target tracking algorithm is used for key monitoring area personnel target tracking; after personnel image information after key monitoring area personnel target detection and tracking, face feature extraction, behavior trajectory feature extraction and skeleton feature extraction are carried out, and feature fusion is carried out; after the features after feature fusion are respectively input into the generated adversarial network and the BiLSTM network which have been trained to complete, energy station abnormal behavior recognition and decision fusion are carried out, and the abnormal behavior category of the key monitoring area of the energy station is obtained; the abnormal behavior includes personnel abnormal wandering behavior, hesitation and retention behavior, violent destruction behavior and illegal operation behavior; according to the abnormal behavior category of the key monitoring area of the energy station, the abnormal behavior level where the abnormal behavior is located is analyzed, and the corresponding degree of behavior warning is carried out; multi-source image acquisition can be carried out through multi-source image acquisition devices, and multi-source image fusion is carried out to establish an image data basis for subsequent abnormal behavior recognition and warning, avoiding low recognition rate of single image data; in addition, target detection and target tracking are carried out on the personnel in the key monitoring area to ensure that personnel can be detected and tracked in the complex scene of the energy station to obtain the position information of each frame of personnel in the video image; the extraction and fusion of face features, behavior trajectory features and skeleton features can more accurately reflect the abnormal behavior of personnel, solve the scene complexity, and it is difficult to reflect the behavior of personnel through a single feature; and the energy station abnormal behavior recognition and decision fusion are carried out in the generated adversarial network and the BiLSTM network, and the abnormal behavior level where the abnormal behavior is located is analyzed, and the corresponding degree of behavior warning is carried out, which can improve the accuracy of abnormal behavior recognition and timely warning, avoid some accidents and economic losses, and ensure the safe and reliable operation of the energy station.

[0047] Other features and advantages will be set forth in the following specification, and in part will be apparent from the description, or can be learned by practice of the application. The objects and other advantages of the application will be realized and attained by the structure particularly pointed out in the written description and claims thereof as well as the appended drawings.

[0048] So that the foregoing aspects and features of the present application can be understood in more detail, a more particular description will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawings and will be described herein below. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0050] Figure 1 A flow chart of an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning is provided.

[0051] Figure 2 A principle schematic block diagram of an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning is provided. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme of the present application will be described clearly and completely below in conjunction with the drawings. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0053] Embodiment 1

[0054] Figure 1 A flow chart of an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning is provided.

[0055] Figure 2 A principle schematic block diagram of an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning is provided.

[0056] As shown in Figure 1 , 2 , the present embodiment 1 provides an energy station abnormal behavior early warning method based on multi-source image recognition and deep learning, which comprises:

[0057] S1, dividing the energy station into a plurality of key monitoring areas, and respectively setting a first image acquisition device, a second image acquisition device and a third image acquisition device in each key monitoring area;

[0058] S2, starting the first image acquisition device to patrol the key monitoring area through scheduling, collecting the image of personnel entering the area, and judging whether personnel enter or not. After judging that personnel enter, the second image acquisition device and the third image acquisition device are started to collect multi-source images.

[0059] S3, image preprocessing, image registration, image fusion and fusion evaluation are performed on the collected multi-source images to obtain multi-source fusion images of the key monitoring area;

[0060] S4, the improved YOLOv5 algorithm is used for key monitoring area personnel target detection on the multi-source fusion images of the key monitoring area, and the DeepSort target tracking algorithm is used for key monitoring area personnel target tracking;

[0061] S5, after the personnel image information after the key monitoring area personnel target detection and tracking, face feature extraction, behavior trajectory feature extraction and skeleton feature extraction are performed, and feature fusion is performed;

[0062] S6, after the feature fusion, the features are input into the trained generative adversarial network and BiLSTM network for energy station abnormal behavior recognition and decision fusion to obtain the abnormal behavior category of the key monitoring area of the energy station; the abnormal behavior includes personnel abnormal wandering behavior, hesitation and retention behavior, violent destruction behavior and illegal operation behavior;

[0063] S7, according to the abnormal behavior category of the key monitoring area of the energy station, the abnormal behavior level of the abnormal behavior is analyzed, and the corresponding degree of behavior warning is performed.

[0064] In this embodiment, the energy station is divided into a plurality of key monitoring areas in S1, including: based on the key equipment in the energy station including heat source unit, heat station, fresh air unit, air conditioning unit, water pump, valve, cooling tower, key pipeline pipe, main electrical equipment, the energy station is divided into a plurality of key monitoring areas;

[0065] The first image acquisition device is arranged at the entrance and exit of each key monitoring area; the second image acquisition device and the third image acquisition device are arranged according to the different shooting angles of the positions near the key equipment; the resolution of the second image acquisition device and the third image acquisition device is higher than that of the first image acquisition device; the types of the second image acquisition device and the third image acquisition device are different, and the number is at least one.

[0066] In this embodiment, in S2, the first image acquisition device is started to patrol the key monitoring area by scheduling to collect personnel entering area images, and to judge whether there is personnel entering. After judging that there is personnel entering, the second image acquisition device and the third image acquisition device are started to collect multi-source images, including:

[0067] The first image acquisition device, the second image acquisition device and the third image acquisition device are classified;

[0068] Firstly, the first image acquisition device is started to patrol each key monitoring area to collect the image of the personnel entering or leaving the area, and to determine whether the personnel enter the key monitoring area. If the personnel enter the key monitoring area, the second image acquisition device and the third image acquisition device corresponding to the key monitoring area are started according to the video image of the personnel entering the area and the position of the detection area, and the video stream image is extracted to collect the multi-source image of the personnel behavior.

[0069] It should be noted that if the first image acquisition device, the second image acquisition device and the third image acquisition device are all turned on and continuously collect video images for a long time, the CPU and the memory will be greatly burdened. Therefore, the second image acquisition device and the third image acquisition device can be turned off first, and the first image acquisition device can be used to determine whether the personnel enter or leave to turn on the second image acquisition device and the third image acquisition device of the corresponding monitoring area. That is, the resources can be dynamically allocated according to the position of the abnormal personnel behavior in the energy station. The number of the second image acquisition device and the third image acquisition device is selected according to the size of the area, the position of the key equipment in the area, the angle, etc.

[0070] In the S3, the image preprocessing includes stretching the gray scale of the image by using gamma correction and suppressing the noise of the image by using a Gaussian curvature filtering algorithm.

[0071] The image registration includes firstly detecting, describing and matching the feature points of the images to be registered by using the scale invariant feature transform (SIFT) algorithm, then deleting the feature vectors with feature values less than a set value in the described feature vectors by using the singular value decomposition (SVD) algorithm to reduce the dimension of the feature vectors, and finally reconstructing the feature point description vectors and performing the initial matching of the feature points by using the descriptor reorganization strategy, and eliminating the error matching points in the matching process to realize the image registration.

[0072] The image fusion includes decomposing the multi-source images into high-frequency and low-frequency images by multi-scale transformation, inputting the high-frequency components and the low-frequency components into a neural network model for fusion to obtain high-low frequency fusion images, and obtaining the fusion image after inverse transformation.

[0073] The fusion evaluation includes subjective evaluation and objective evaluation. The subjective evaluation is to directly evaluate the overall fusion image by observing the image texture details, clarity and overall information amount by the human eye. The objective evaluation is to evaluate based on the information entropy, standard deviation, average gradient, mutual information, structural similarity and edge information retention value indexes of the image.

[0074] In the embodiment, the S4, the key monitoring area multi-source fusion image is detected by using the improved YOLOv5 algorithm, including:

[0075] The YOLOv5 algorithm is improved: in the network structure of the original YOLOv5 model, CBAM attention mechanism modules are added at the Backbone backbone network, the Neck unit and the Prediction unit, at the same time, the Ghost lightweight convolution layer is introduced to replace the general convolution layer in the Backbone backbone network, the GhostBottleneck structure is introduced to replace the CSP structure in the Backbone, and the weighted bidirectional feature pyramid is introduced to replace the bottom-up feature pyramid in the Neck unit;

[0076] The key monitoring area multi-source fusion image is labeled to obtain a labeled data set, and the improved YOLOv5 algorithm model is trained to obtain a personnel target detection model;

[0077] The key monitoring area multi-source fusion image to be detected is input into the personnel target detection model, the personnel and position in the key monitoring area are detected, and the positioning frame of the personnel is output.

[0078] In the embodiment, the key monitoring area personnel target tracking is performed by using the DeepSort target tracking algorithm, including:

[0079] The position and state of the personnel target in the next frame are predicted by using the Kalman filter algorithm;

[0080] The Hungarian algorithm is used for matching to obtain the trajectory of the personnel before and after the video image, and the Mahalanobis distance between the personnel detection frame and the personnel tracking frame is calculated, when the distance is less than a set threshold, the two are associated with each other, and the matching is successful; otherwise, the personnel position is recalculated.

[0081] The tracking frame parameters that have been matched are updated by using the Kalman filter update formula, and the personnel target at the next moment is predicted; when the result predicted by the updated parameters cannot be matched, it is indicated that the key monitoring area personnel target has been lost, and the tracking frame is deleted.

[0082] In the embodiment, the S5, after the personnel image information after the key monitoring area personnel target detection and tracking is extracted, the feature fusion is performed, including:

[0083] The personnel image information after personnel target detection and tracking in the key monitoring area is input into a pre-trained graph convolutional neural network model to extract face features, and the face information in the database is compared. If the comparison is successful, it indicates that the personnel is a normal operation personnel; otherwise, it indicates that the personnel is an abnormal personnel;

[0084] A trajectory model including personnel position features, direction features and speed features is constructed. After the position of each trajectory point is extracted from the personnel image information, the differential data in the position direction is calculated, and the direction and speed information of the personnel target are represented by the differential data. The position features, direction features and speed are input into a pre-trained graph convolutional neural network model to extract behavior trajectory features;

[0085] A skeleton sequence is extracted from the personnel image by using a skeleton extraction algorithm OpenPose, and a pre-trained graph convolutional neural network model is used to extract a spatio-temporal skeleton graph G(V, E) to obtain skeleton features. V is a personnel skeleton joint node, and E is an edge of the skeleton graph. The skeleton sequence includes N joint nodes in each frame;

[0086] The extracted face features, behavior trajectory features and skeleton features are fused to form a face-trail-skeleton fusion feature.

[0087] In actual applications, there are many kinds of abnormal behaviors of energy stations, such as abnormal wandering behavior, hesitation and retention behavior, violent destruction behavior and illegal operation behavior, etc. These abnormal behaviors have some characteristics, and are sudden, with short process duration. Some information of abnormal behaviors can be obtained through joint node flow information, skeleton flow information and skeleton motion flow information. The behavior trajectory features also include discrete curvature entropy, motion distance and displacement. The discrete curvature entropy is used to represent the degree of curve bending, and describes the degree of disorder of the personnel trajectory. The entropy value is represented as H i (dir) is the histogram of the i-th trajectory motion direction change value. Assuming that the personnel moving frame is k frames, the position point information is (x, y), and the displacement size in the k frames is The distance is obtained by accumulating and summing all displacements; the speed change can be expressed by the second difference d 2 x=x i -2x i-1 +x i-2 , d 2 y=y i -2y i-1 +y i-2 ;

[0088] It should be noted that because the video image information is relatively complex, it is difficult to reflect the behavior of the personnel through a single feature, and therefore the extraction and fusion of the face feature, the behavior trajectory feature and the skeleton feature can more accurately reflect the abnormal behavior of the personnel. In addition, rigid features can also be extracted, including whether the clothing of the personnel is a work uniform and whether the head is wearing a safety helmet.

[0089] In the embodiment, the graph convolutional neural network is an attention-enhanced graph convolutional neural network model, which includes one network full connection layer, one long short-term memory network, three attention-enhanced graph convolutional long short-term memory networks and two full connection layers; in the training process of the graph convolutional neural network model, the fused features are taken as input data of the model, the cross-entropy loss and the attention regularization loss of the network are calculated through each layer of the model, and the network parameters are updated using back propagation, and the training is ended when the training number reaches a set value or the network loss value is lower than a preset threshold.

[0090] In the embodiment, the S6, the features after the feature fusion are respectively input into the trained generative adversarial network and the BiLSTM network for energy station abnormal behavior recognition and decision fusion, and the abnormal behavior category of the key monitoring area of the energy station is obtained, including:

[0091] The features after the feature fusion are input into the trained generative adversarial network for energy station abnormal behavior recognition, and an energy station abnormal behavior recognition result I is obtained;

[0092] The features after the feature fusion are input into the trained BiLSTM network for energy station abnormal behavior recognition, and an energy station abnormal behavior recognition result II is obtained;

[0093] After decision fusion of the energy station abnormal behavior recognition I and the energy station abnormal behavior recognition result II, the abnormal behavior category of the key monitoring area of the energy station is obtained;

[0094] The generative adversarial network includes a generator network, a discriminator network and an optical flow estimation network.

[0095] The BiLSTM network includes a BiLSTM neural network layer and an output layer, wherein the BiLSTM neural network layer includes an input layer, a forward propagation layer and a backward propagation layer, the number of nodes of the input layer, the forward propagation layer, the backward propagation layer and the output layer is set first, and the weights between the nodes of adjacent layers are randomly set.

[0096] It should be noted that the generator network is composed of the spatio-temporal feature encoder and the prediction branch in the spatio-temporal feature decoder; the first N-1 frame continuous video image sequence is used as the input of the generator network, the spatio-temporal features of the sequence are extracted by the spatio-temporal feature encoder, and the prediction branch in the spatio-temporal feature decoder is used for the prediction of the Nth frame image; the image loss is calculated for the predicted image and the real image; the predicted image and the real image are input into the discriminator network to calculate the adversarial loss; the real N-1th frame image is input into the optical flow estimation network together with the predicted Nth frame image and the real Nth frame image to obtain the predicted optical flow information and the real optical flow information, and the optical flow loss is calculated; the total loss is calculated based on the image loss, the adversarial loss and the optical flow loss, and the generator network and the optical flow estimation network are trained; the predicted Nth frame image is obtained by using the generator network for prediction, and the peak signal-to-noise ratio is calculated with the real Nth frame image; the peak signal-to-noise ratio is compared with the set threshold value to identify the abnormal behavior.

[0097] The training process of the BiLSTM neural network: input the training set into the input layer of the BiLSTM neural network, learn the parameters of each neural network layer; finally, input the test set into the BiLSTM neural network to identify abnormal behavior.

[0098] Before identifying abnormal behavior, if the training data is less, the image data of abnormal behavior can be expanded by using the generative adversarial network: using the data enhancement technology based on the generative adversarial network, synthetic artificial image data similar to the original image data distribution and maintaining the dynamic characteristics of the original image data, correcting the sample imbalance problem, the input of the generator in the generative adversarial network is random noise, and the output is the synthesized abnormal behavior data; the input of the discriminator is the behavior data of a certain category in the original data set and the output sample of the generator, and the output is the judgment of the generated sample; when the discriminator cannot distinguish the original data from the synthetic data, the generative adversarial network converges, at this time the generator can synthesize realistic data, and the training of the entire generative adversarial network is optimized; use the trained generative adversarial network to expand the image data of different abnormal behaviors, and form a new data set with the original image data.

[0099] In the embodiment, the S7 analyzes the abnormal behavior level in which the abnormal behavior is located according to the abnormal behavior category of the key monitoring area of the energy station, and performs behavior warning of a corresponding degree, including: matching the abnormal behavior and the pre-set abnormal behavior level according to the abnormal behavior category of the key monitoring area of the energy station to obtain the abnormal behavior level in which the abnormal behavior is located, if it is a first-level abnormal behavior, sound and light alarm should be performed, and the superior department should be reported immediately, the running state and parameters of the monitoring equipment should be monitored and remote control should be taken; if it is a second-level abnormal behavior, sound and light alarm should be performed, and someone should be sent to the area to check; if it is a third-level abnormal behavior, sound and light alarm should be performed.

[0100] In several embodiments provided in the present application, it should be understood that the disclosed system and method can also be implemented in other manners. The above described system embodiments are merely exemplary. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operation of the system, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment or a portion of code which comprises one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the accompanying drawings. For example, two blocks shown in succession can in fact be executed substantially concurrently or in the reverse order, depending upon the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by dedicated hardware-based systems which perform the specified functions or acts, or can be implemented by a combination of dedicated hardware and computer instructions.

[0101] In addition, the various functional modules in the embodiments of the present application can be integrated into one independent part, or each functional module can exist alone, or two or more functional modules can be integrated into one independent part. When the functions are implemented in the form of software functional modules and sold or used as an independent product, they can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially, or the part that contributes to the prior art, or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the methods of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various other media that can store program codes.

[0102] Based on the above ideal embodiments according to the present application, through the above description, relevant personnel can make various changes and modifications without deviating from the scope of the technical idea of the present application. The technical scope of the present application is not limited to the content in the specification, and must be determined according to the scope of the claims.

Claims

1. A multi-source image recognition and deep learning-based energy station abnormal behavior early warning method, characterized in that, The method comprises the following steps: S1, dividing the energy station into a plurality of key monitoring areas, and respectively arranging a first image acquisition device, a second image acquisition device and a third image acquisition device in each key monitoring area; S2, starting the first image acquisition device to patrol the key monitoring area through scheduling, collecting the image of personnel entering the area, and judging whether personnel enter or not, and after judging that personnel enter, starting the second image acquisition device and the third image acquisition device to collect multi-source images; S3, performing image preprocessing, image registration, image fusion and fusion evaluation on the collected multi-source images to obtain a multi-source fusion image of the key monitoring area; S4, performing key monitoring area personnel target detection on the multi-source fusion image of the key monitoring area by using an improved YOLOv5 algorithm, and performing key monitoring area personnel target tracking by using a DeepSort target tracking algorithm; The key monitoring area personnel target detection on the multi-source fusion image of the key monitoring area by using the improved YOLOv5 algorithm comprises: Improving the YOLOv5 algorithm: adding a CBAM attention mechanism module at the Backbone main network, the Neck unit and the Prediction unit in the original YOLOv5 model network structure, simultaneously introducing a Ghost lightweight convolutional layer to replace a general convolutional layer in the Backbone main network, introducing a GhostBottleneck structure to replace a CSP structure in the Backbone, and introducing a weighted bidirectional feature pyramid to replace a bottom-up feature pyramid in the Neck unit; Labeling the multi-source fusion image of the key monitoring area to obtain a labeled data set, and after training the improved YOLOv5 algorithm model, obtaining a personnel target detection model; Inputting the multi-source fusion image to be detected of the key monitoring area into the personnel target detection model to detect the personnel and the position of the key monitoring area, and outputting the positioning frame of the personnel; S5, after performing face feature extraction, behavior trajectory feature extraction and skeleton feature extraction on the personnel image information after the key monitoring area personnel target detection and tracking, and performing feature fusion; S6, after inputting the features after the feature fusion into the trained generative adversarial network and BiLSTM network, performing energy station abnormal behavior recognition and decision fusion, and obtaining the abnormal behavior category of the key monitoring area of the energy station; the abnormal behavior comprises personnel abnormal wandering behavior, hesitation and retention behavior, violent destruction behavior and illegal operation behavior; S7, according to the abnormal behavior category of the key monitoring area of the energy station, analyzing the abnormal behavior level in which the abnormal behavior is located, and performing behavior warning in a corresponding degree.

2. The energy plant abnormal behavior early warning method of claim 1, wherein, In S1, the energy station is divided into a plurality of key monitoring areas, which comprises: based on the key equipment in the energy station comprising a heat source unit, a heat station, a fresh air unit, an air conditioning unit, a water pump, a valve, a cooling tower, a key pipeline pipe and main electrical equipment, the energy station is divided into a plurality of key monitoring areas; The first image acquisition device is arranged at the entrance and exit of each key monitoring area; the second image acquisition device and the third image acquisition device are arranged according to different shooting angles near the key equipment; the resolution of the second image acquisition device and the third image acquisition device is higher than that of the first image acquisition device; the types of the second image acquisition device and the third image acquisition device are different, and the number is at least one.

3. The energy plant abnormal behavior early warning method of claim 1, wherein, S2, by scheduling, starting the first image acquisition device to patrol the key monitoring area, collecting the image of the personnel entering the area, judging whether there is personnel entering, after judging that there is personnel entering, scheduling starting the second image acquisition device and the third image acquisition device to collect multi-source images, including: Classifying the first image acquisition device, the second image acquisition device and the third image acquisition device; First, by scheduling, starting the first image acquisition device to patrol each key monitoring area, collecting the image of the personnel entering and exiting the area, judging whether there is personnel entering the key monitoring area, if there is personnel entering, according to the video image of the personnel entering the area collected, calculating the direction of the personnel advancing, and then according to the position of the detected area where the personnel enter, scheduling starting the second image acquisition device and the third image acquisition device corresponding to the key monitoring area, extracting the video stream image, and collecting multi-source images of personnel behavior.

4. The energy plant abnormal behavior early warning method of claim 1, wherein, In S3, the image preprocessing includes stretching the gray scale of the image by using gamma correction and suppressing the noise of the image by using Gaussian curvature filtering algorithm; The image registration includes first detecting, describing and matching feature points of the images to be registered by using scale invariant feature transform (SIFT) algorithm, then deleting the feature vectors with feature values less than a set value in the described feature vectors by using singular value decomposition (SVD) algorithm, reducing the dimension of the feature vectors, finally reconstructing the feature point description vector and performing initial matching of the feature points by using the descriptor reorganization strategy, and removing the wrong matching points in the matching process to realize image registration; The image fusion includes decomposing the multi-source images into high-frequency and low-frequency images by multi-scale transformation, inputting the high-frequency and low-frequency components into a neural network model for fusion to obtain high-low frequency fusion images, and then performing inverse transformation to obtain the fusion image; The fusion evaluation includes subjective evaluation and objective evaluation; the subjective evaluation is to directly evaluate the overall fusion image by observing the image texture details, clarity and overall information amount with the human eye; the objective evaluation is based on the information entropy, standard deviation, average gradient, mutual information, structural similarity and edge information retention value indexes of the image.

5. The energy plant abnormal behavior early warning method of claim 1, wherein, The key monitoring area personnel target tracking by using DeepSort target tracking algorithm includes: Using Kalman filtering algorithm to predict the position and state of the personnel target in the next frame; Using Hungarian algorithm for matching to obtain the trajectory of the personnel before and after the video image, and calculating the Mahalanobis distance between the personnel detection frame and the personnel tracking frame, when the distance is less than a set threshold, the two are associated with each other, and the matching is successful; otherwise, the personnel position is recalculated; The tracking frame parameters that have been matched are updated using Kalman filter update formula, and the personnel target at the next time is predicted; when the result predicted by the updated parameters cannot be matched, it is indicated that the personnel target in the current key monitoring area has been lost, and the tracking frame is deleted.

6. The energy plant abnormal behavior early warning method of claim 1, wherein, The personnel image information after the personnel target detection and tracking in the key monitoring area is subjected to face feature extraction, behavior trajectory feature extraction and skeleton feature extraction, and then feature fusion is performed, including: The personnel image information after the personnel target detection and tracking in the key monitoring area is input into a pre-trained graph convolutional neural network model for face feature extraction, and the face information in the database is compared; if the comparison is successful, it is indicated that the personnel is a normal worker; otherwise, it is indicated that the personnel is an abnormal personnel; A trajectory model including personnel position features, direction features and speed features is constructed; after the position of each trajectory point is extracted from the personnel image information, the differential data in the position direction is calculated, and the direction and speed information of the personnel target are represented by the differential data; the position features, direction features and speed are input into a pre-trained graph convolutional neural network model for behavior trajectory feature extraction; A skeleton sequence is extracted from the personnel image by using a skeleton extraction algorithm OpenPose, and a pre-trained graph convolutional neural network model is used to extract a spatiotemporal skeleton graph G(V, E) to obtain skeleton features; V is a personnel skeleton joint node; E is an edge of the skeleton graph; the skeleton sequence includes N joint nodes in each frame; The extracted face features, behavior trajectory features and skeleton features are subjected to feature fusion to form a fusion feature of face-trail-skeleton.

7. The energy plant abnormal behavior early warning method of claim 6, wherein, The graph convolutional neural network is an attention-enhanced graph convolutional neural network model, including one network full connection layer, one long short-term memory network, three attention-enhanced graph convolutional long short-term memory networks and two full connection layers; during the training of the graph convolutional neural network model, the fusion features are used as input data of the model, the cross-entropy loss and attention regularization loss of the network are calculated through each layer of the model, and the network parameters are updated using back propagation; the training is ended when the number of training times reaches a set value or the network loss value is lower than a pre-set threshold.

8. The energy plant abnormal behavior early warning method of claim 1, wherein, The S6 includes: The features after the feature fusion are input into a pre-trained generative adversarial network and a BiLSTM network for energy station abnormal behavior recognition and decision fusion to obtain an abnormal behavior category of the key monitoring area of the energy station. The features after the feature fusion are input into a pre-trained generative adversarial network for energy station abnormal behavior recognition to obtain an energy station abnormal behavior recognition result I. The features after the feature fusion are input into a pre-trained BiLSTM network for energy station abnormal behavior recognition to obtain an energy station abnormal behavior recognition result II. The energy station abnormal behavior recognition I and the energy station abnormal behavior recognition result II are subjected to decision fusion to obtain an abnormal behavior category of the key monitoring area of the energy station. The generative adversarial network includes a generator network, a discriminator network and an optical flow estimation network. The BiLSTM network comprises a BiLSTM neural network layer and an output layer, wherein the BiLSTM neural network layer comprises an input layer, a forward propagation layer and a backward propagation layer, the number of nodes of the input layer, the forward propagation layer, the backward propagation layer and the output layer is set in advance, and the weights between nodes of adjacent layers are randomly set.

9. The energy plant abnormal behavior early warning method of claim 1, wherein, The S7 analyzes the abnormal behavior level in which the abnormal behavior is located according to the abnormal behavior category of the key monitoring area of the energy station, and performs a corresponding degree of behavior warning, comprising: according to the abnormal behavior category of the key monitoring area of the energy station, matching the abnormal behavior with the pre-set abnormal behavior level to obtain the abnormal behavior level in which the abnormal behavior is located, if it is a first-level abnormal behavior, a sound and light alarm should be performed, and immediately reported to the superior department, the running state and parameters of the monitoring equipment are monitored, and remote control is adopted; if it is a second-level abnormal behavior, a sound and light alarm should be performed, and a person should be sent to the area to check; if it is a third-level abnormal behavior, a sound and light alarm should be performed.

Citation Information

Patent Citations

  • Plant personnel abnormal behavior detection method based on deep learning

    CN112597877A

  • Personnel abnormal behavior early warning method and system based on video data machine learning, and computer equipment

    CN113850229A