Transformer substation behavior detection method and system based on video stream analysis and processor

By performing frame extraction, target detection, and image feature extraction on the video stream of substation workers, and combining the Transformer encoder and RestNet50 model for behavior analysis, abnormal warning information is generated. This solves the problem of inaccurate detection of violations by substation workers and achieves accurate violation identification and timely alarm.

CN120894818APending Publication Date: 2025-11-04STATE GRID ANHUI ELECTRIC POWER CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510870309.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

The existing intelligent judgment model for unplanned operations in substations has problems with false alarms from equipment and inaccurate identification of the nighttime inspection environment, resulting in inaccurate detection of violations by staff.

Method used

By acquiring video streams from substation staff, frame extraction, target detection, and image feature extraction are performed. Behavioral analysis is then conducted using a Transformer encoder and a RestNet50 model to generate abnormal warning messages, which are then pushed to IoT sensing devices.

Benefits of technology

It enables accurate identification and timely alarm of violations by substation staff, reduces interference from manual verification, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894818A_ABST
    Figure CN120894818A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a transformer substation behavior detection method and system based on video stream analysis and a processor, and belongs to the technical field of transformer substation video detection. The transformer substation behavior detection method comprises the following steps: acquiring a video stream of a transformer substation about the action of a worker; performing frame extraction processing on the video stream to obtain an image containing a worker; performing target detection on the image to obtain image features about the staff; performing action analysis on the image features, and performing behavior detection to generate abnormal warning information; and prompting the abnormal warning information about the worker and the corresponding video clip through the Internet of Things sensing equipment or information so as to realize the pushing of the alarm. The transformer substation behavior detection method can accurately identify the violation condition of the staff in the transformer substation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of substation video detection technology, and more specifically to a substation behavior detection method, system, and processor based on video stream analysis. Background Technology

[0002] Currently, intelligent analysis models for unplanned operations in substations still suffer from false alarms due to issues such as substation equipment and nighttime inspection environments. This is primarily due to inaccurate identification of inspection personnel. Manual verification by supervisory center staff to determine if unplanned operations are causing interference is crucial. Only with accurate identification can the behavior of substation personnel be detected, enabling timely alerts when violations are discovered. Therefore, a substation behavior detection method based on video stream analysis is needed to accurately identify violations by substation personnel. Summary of the Invention

[0003] The purpose of this invention is to provide a substation behavior detection method, system, and processor based on video stream analysis. This substation behavior detection method can accurately identify violations by staff in substations.

[0004] To achieve the above objectives, embodiments of the present invention provide a substation behavior detection method based on video stream analysis, the substation behavior detection method comprising: Obtain video streams from the substation showing the actions of staff; The video stream is subjected to frame extraction to obtain images containing staff members; Target detection is performed on the image to obtain image features related to the staff; The image features are subjected to motion analysis and behavior detection to generate abnormal warning information; The system will push alerts by sending notifications of abnormal warnings about staff and corresponding video clips through IoT sensing devices or information systems.

[0005] Optionally, the video stream is subjected to frame extraction to obtain an image containing the staff, including: The video stream is acquired, and frames are extracted from the video stream at a fixed frame extraction frequency; Determine whether the extracted frame image contains staff members; If the image obtained by frame extraction contains workers, increase the frame extraction frequency; If the image obtained by frame extraction does not contain workers, reduce the frame extraction frequency; Collect all the images containing staff obtained from frame extraction.

[0006] Optionally, target detection is performed on the image to obtain image features about the staff, including: Acquire an image containing staff members, and perform background segmentation on the image to obtain a filtered foreground image containing staff members; Based on the acquired foreground image, edge optimization is performed on the objects in the foreground image to obtain an optimized foreground image; Determine whether any objects are missing in the optimized foreground image; If any parts are missing, repair them. The repaired foreground image and the foreground image without missing parts are subjected to feature transformation to obtain the image features of the foreground image; Obtain the second feature of the image containing the worker from different frames corresponding to the image feature location information in the foreground image; The image features and second features of the foreground image are fused to obtain image features about the staff.

[0007] Optionally, an image containing staff is acquired, and the image is segmented to obtain a filtered foreground image containing staff, including: Obtain an image containing the staff member, and feed the image into RestNet50 for feature extraction and straighten the extracted features; The straightened features are combined with the positional encoding to obtain the sequence features, which are then input into the Transformer encoder. The Transformer encoder learns global information based on the input sequence features and models the global context through a self-attention mechanism. Then, the Transformer decoder generates prediction boxes. The best match between the predicted bounding box and the true value obtained by formula (1): Formula (1), in, This represents the optimal permutation. Indicates the number of prediction boxes. This represents a set of predicted bounding boxes and actual values. Indicates the label, Represents the actual value. Indicates the prediction box. This represents the cost of matching a single true value with a predicted bounding box. The loss function is calculated to find the best match between the predicted bounding box and the ground truth, and the parameters in the RestNet50, Transformer encoder, and Transformer decoder are adjusted based on minimizing the loss function. The predicted bounding boxes are obtained based on the adjusted parameters. The confidence scores of the predicted bounding boxes are calculated, and the predicted bounding boxes with confidence scores greater than a preset threshold are selected as the foreground images.

[0008] Optionally, based on the acquired foreground image, edge optimization is performed on the objects in the foreground image to obtain an optimized foreground image, including: Acquire a foreground image containing the staff, and locate the pixel outline boundaries in the foreground image; Based on the obtained pixel outline boundary, select the pixel points of the pixel outline boundary, and with the selected pixel point as the center, obtain all pixel points within the 5*5 neighborhood. Based on the obtained pixels, the initial weight of each pixel is obtained using formula (2): Formula (2), in, Represents pixels The initial weights, The pixel representing the center point, Indicates the weighting parameter; The initial weights of each obtained pixel are normalized, and the pixel value of each pixel is multiplied by the normalized initial weights to update the pixel value of the center pixel. The pixels of all pixel boundaries in the foreground image are traversed to perform edge optimization on the objects in the foreground image, so as to obtain the optimized foreground image.

[0009] Optionally, the foreground image undergoes feature transformation to obtain image features of the foreground image, including: The foreground image is acquired, and features are extracted from the foreground image to obtain feature points related to the foreground image; Multiple composite geometric transformations are performed on the foreground image to obtain foreground images of various shapes and angles; Features of the foreground image after the composite geometric transformation are extracted to obtain new feature points; The new feature points are mapped back to the feature points of the initial foreground image to obtain aggregated feature points, and the image features of the foreground image are obtained based on the aggregated feature points.

[0010] Optionally, the image features and the second feature of the foreground image are fused to obtain image features about the worker, including: Select three feature points from the image features in the foreground image; Select three feature points from the second feature corresponding to the three feature points of the image features in the foreground image; The region where the line connecting the three corresponding feature points in the second feature overlaps with the region where the line connecting the three feature points in the foreground image is the matching feature region. Based on the image features of the foreground image, the matching feature region is fused into the image features of the foreground image to obtain image features about the staff.

[0011] On the other hand, the present invention can also provide a substation behavior detection system based on video stream analysis, the substation behavior detection system comprising: The facial video acquisition module is used to acquire video streams of staff members; The detection module is used to execute the substation behavior detection method based on video stream analysis as described above, according to the acquired video stream.

[0012] On the other hand, the present invention may also provide a processor for running a program, wherein the program is run to execute: the substation behavior detection method based on video stream analysis as described above.

[0013] Through the above technical solution, the substation behavior detection method, system, and processor based on video stream analysis provided by this invention acquires video streams of substation workers' actions, and then performs frame extraction processing on the video streams to obtain images containing workers. After obtaining the images, target detection can be performed on the obtained images to obtain image features related to workers. After obtaining the image features, action analysis and behavior detection can be performed on the acquired image features, thereby generating abnormal alarm information when violations are detected in the images. After receiving the abnormal alarm information, the abnormal warning information about workers and the corresponding video clips can be sent through IoT sensing devices or information systems to achieve alarm push. This substation behavior detection method can accurately identify violations by workers in substations.

[0014] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a substation behavior detection method based on video stream analysis according to an embodiment of the present invention; Figure 2This is a flowchart of frame extraction for a substation behavior detection method based on video stream analysis according to an embodiment of the present invention. Figure 3 This is a flowchart illustrating the acquisition of image features in a substation behavior detection method based on video stream analysis according to an embodiment of the present invention. Figure 4 This is a flowchart of a substation behavior detection method based on video stream analysis according to an embodiment of the present invention, which involves acquiring a foreground image. Figure 5 This is a flowchart of the foreground image optimization of a substation behavior detection method based on video stream analysis according to an embodiment of the present invention; Figure 6 This is a flowchart of the feature transformation of a substation behavior detection method based on video stream analysis according to an embodiment of the present invention; Figure 7 This is a flowchart illustrating the feature fusion of a substation behavior detection method based on video stream analysis according to an embodiment of the present invention. Detailed Implementation

[0016] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0017] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0018] Figure 1 This is a flowchart of a substation behavior detection method based on video stream analysis according to an embodiment of the present invention. In this invention, the flow of the detection method may include: In step S1, a video stream of the substation's staff actions is acquired.

[0019] In step S2, the video stream is subjected to frame extraction to obtain an image containing the staff.

[0020] In step S3, target detection is performed on the image to obtain image features related to the staff.

[0021] In step S4, motion analysis and behavior detection are performed on the image features to generate abnormal warning information.

[0022] In step S5, abnormal warning information about staff and corresponding video clips are sent through IoT sensing devices or information to push alarms.

[0023] In this invention, when performing substation behavior detection, a video stream of the substation's workers' actions can be acquired first. After obtaining the video stream, frame extraction can be performed to obtain images containing the workers. After obtaining the images, object detection can be performed to obtain image features related to the workers. After obtaining the image features, action analysis and behavior detection can be performed on the acquired image features. In this application, a parallel point detection and point matching (PPDM) framework for human interaction detection can be used for behavior detection. In PPDM, HOI is defined as a point triple <human point, interaction point, object point>. The human and object points are the centers of the detection boxes, and the interaction point is the center point of the line connecting the human and object points. PPDM contains two parallel branches: a point detection branch and a point matching branch. The point detection branch predicts three points. Simultaneously, the point matching branch predicts two displacements from the interaction point to the corresponding human and object points. Human points and object points from the same interaction point are considered as matching pairs. In the parallel architecture, interaction points implicitly provide context and regularization for human and object detection, suppressing isolated detection boxes of meaningless HOIs and improving the accuracy of HOI detection. Furthermore, matching between human and object detection boxes is only applicable to a limited number of filtered candidate interaction points, thus saving significant computational costs.

[0024] This allows for the generation of abnormal alarm information when violations are detected in images. Upon receiving the alarm, the abnormal warning information regarding the staff member and the corresponding video clip can be sent via IoT sensing devices or information systems, thus enabling alarm push notifications. This substation behavior detection method can accurately identify violations by staff members in substations.

[0025] In one embodiment of the present invention, such as Figure 2 As shown, the frame extraction process may include: In step S6, the video stream is acquired, and frames are extracted from the video stream using a fixed frame extraction frequency.

[0026] In step S7, it is determined whether the image obtained by frame extraction contains staff members.

[0027] In step S8, if the image obtained by frame extraction contains workers, the frame extraction frequency is increased.

[0028] In step S9, if the image obtained by frame extraction does not contain staff members, the frame extraction frequency is reduced.

[0029] In step S10, all the images containing staff obtained from frame extraction are collected.

[0030] In this invention, when performing frame extraction on a video stream, the video stream can be acquired first, and then frames can be extracted from the acquired video stream at a fixed extraction frequency. After frame extraction, it can be determined whether the extracted images contain staff members. If the extracted images contain staff members, the extraction frequency can be increased. If the extracted images do not contain staff members, the extraction frequency can be decreased. This method can effectively extract frames from the video stream that contain staff members. After frame extraction, all extracted images containing staff members can be aggregated.

[0031] In one embodiment of the present invention, such as Figure 3 As shown, the process of obtaining image features may include: In step S11, an image containing staff is acquired, and the image is segmented to obtain a filtered foreground image containing staff.

[0032] In step S12, the edges of the objects in the foreground image are optimized based on the obtained foreground image to obtain an optimized foreground image.

[0033] In step S13, it is determined whether there are any missing objects in the optimized foreground image.

[0034] In step S14, if there are any missing parts, the missing parts are repaired.

[0035] In step S15, feature transformation is performed on the repaired foreground image and the foreground image without missing features to obtain the image features of the foreground image.

[0036] In step S16, the second feature of the image containing the worker in a different frame corresponding to the image feature location information in the foreground image is obtained.

[0037] In step S17, the image features and the second feature of the foreground image are fused to obtain the image features of the worker.

[0038] In this invention, when acquiring image features, an image containing staff can be acquired first, and background segmentation can be performed on the obtained image to obtain a filtered foreground image containing staff. After obtaining the foreground image, edge optimization can be performed on the objects in the foreground image to obtain an optimized foreground image. After obtaining the optimized foreground image, it can be determined whether there are any missing objects in the optimized foreground image. If missing objects are found, they can be repaired. This repair can be performed using the Kalman filter algorithm. After the repair is completed, feature transformation can be performed on the repaired foreground image and the foreground image without missing objects to obtain the image features of the foreground image. These image features can be refined image features to facilitate more accurate recognition fusion in the future. After obtaining the image features in the foreground image, second features of images containing staff in different frames corresponding to the image feature position information in the foreground image can be obtained. After obtaining the second features, the second features can be fused with the image features of the foreground image to obtain more refined image features about the staff.

[0039] In one embodiment of the present invention, such as Figure 4 As shown, the process of acquiring the foreground image may include: In step S18, an image containing staff members is acquired, and the image is fed into RestNet50 for feature extraction, and the extracted features are straightened.

[0040] In step S19, the straightened features are added to the position code to obtain the sequence features, which are then input into the Transformer encoder.

[0041] In step S20, the Transformer encoder learns global information based on the input sequence features and models the global context through a self-attention mechanism. Then, the Transformer decoder generates prediction boxes.

[0042] In step S21, the best match between the predicted bounding box and the true value is obtained through formula (1): Formula (1), in, This represents the optimal permutation. Indicates the number of prediction boxes. This represents a set of predicted bounding boxes and actual values. Indicates the label, Represents the actual value. Indicates the prediction box. This represents the cost of matching a single true value with a predicted bounding box.

[0043] In step S22, the parameters in RestNet50, Transformer encoder, and Transformer decoder are adjusted according to the loss function that calculates the best match between the predicted bounding box and the ground truth value, and the parameters are adjusted according to minimizing the loss function.

[0044] In step S23, a prediction box is obtained based on the adjusted parameters, the confidence level of the prediction box is calculated, and the prediction box with a confidence level greater than a preset threshold is obtained as the foreground image.

[0045] In this invention, after acquiring an image containing staff, the image can be fed into RestNet50 to extract features from the image. The extracted features can be straightened to facilitate better subsequent feature learning. After straightening, the straightened features can be added to the positional encoding to obtain positional encoded sequence features. After obtaining the sequence features, the sequence features are input into the Transformer encoder. The Transformer encoder can learn global information based on the input sequence features and can model the global context through a self-attention mechanism. Then, the Transformer decoder can generate prediction boxes. After generating the prediction boxes, the best match between the prediction boxes and the ground truth can be obtained using formula (1). Based on the obtained best match, the loss function of the best match between the prediction boxes and the ground truth can be calculated. The parameters in RestNet50, Transformer encoder, and Transformer decoder can be adjusted according to minimizing the loss function to make the obtained prediction boxes more accurate. For the prediction boxes obtained after adjusting the parameters, the confidence of the prediction boxes is calculated, and prediction boxes with a confidence greater than a preset threshold can be obtained as foreground images. The prediction bounding box can include information such as staff members; generally, the box containing staff members or similar objects is used as the foreground. Therefore, the foreground image can be obtained through this prediction bounding box.

[0046] In one embodiment of the present invention, such as Figure 5 As shown, the process for optimizing the foreground image may include: In step S24, a foreground image containing the staff is acquired, and the pixel outline boundaries in the foreground image are located.

[0047] In step S25, based on the obtained pixel outline, the pixels of the pixel outline are selected, and with the selected pixels as the center, all pixels in the 5*5 neighborhood are obtained.

[0048] In step S26, based on the obtained pixel points, the initial weight of each pixel point is obtained using formula (2): Formula (2), in, Represents pixels The initial weights, The pixel representing the center point, This represents the weighting parameter.

[0049] In step S27, the initial weight of each obtained pixel is normalized, and the result of multiplying the pixel value of each pixel by the normalized initial weight is accumulated to update the pixel value of the center pixel.

[0050] In step S28, all pixels of the pixel outline in the foreground image are traversed to perform edge optimization on the objects in the foreground image to obtain an optimized foreground image.

[0051] In this invention, when performing edge optimization on objects in a foreground image, a foreground image containing workers can be obtained first, and pixel contour boundaries in the foreground image can be found. After obtaining the pixel contour boundaries, pixels within the pixel contour boundaries can be selected, and then all pixels within a 5x5 neighborhood can be obtained with the selected pixels as the center. After obtaining the pixels within the 5x5 neighborhood, the initial weight of each pixel can be obtained using formula (2). By normalizing the initial weights of each pixel within the obtained 5x5 neighborhood, the result of multiplying the pixel value of each pixel by the normalized initial weights can be accumulated, thereby updating the pixel value of the central pixel. This operation makes the influence of pixels farther away from the central pixel within the 5x5 neighborhood smaller, and the influence of pixels closer to the central pixel larger, so as to obtain the required pixel contour boundaries. After traversing all the pixels within the pixel contour boundaries in the foreground image, edge optimization can be performed on the objects in the foreground image, thereby obtaining the optimized foreground image.

[0052] In one embodiment of the present invention, such as Figure 6 As shown, the feature transformation process can include: In step S29, a foreground image is acquired, and feature extraction is performed on the foreground image to obtain feature points related to the foreground image.

[0053] In step S30, the foreground image undergoes multiple composite geometric transformations to obtain foreground images of various shapes and angles.

[0054] In step S31, features of the foreground image after composite geometric transformation are extracted to obtain new feature points.

[0055] In step S32, the new feature points are mapped back to the feature points of the initial foreground image to obtain aggregated feature points, and the image features of the foreground image are obtained based on the aggregated feature points.

[0056] In this invention, when performing feature transformation on a foreground image, the foreground image can be acquired first, and then features can be extracted from it to obtain feature points. Multiple composite geometric transformations can be performed on the foreground image to obtain foreground images of various shapes and angles. After the composite geometric transformation, features of the transformed foreground image can be extracted to obtain new feature points. These new feature points are then mapped back to the feature points of the initial foreground image to obtain aggregated feature points. Based on these aggregated feature points, the image features of the foreground image can be obtained. Because of the aggregation of these feature points, the image features are more accurate.

[0057] In one embodiment of the present invention, such as Figure 7 As shown, the feature fusion process may include: In step S33, three feature points of the image features in the foreground image are selected.

[0058] In step S34, three feature points from the second feature corresponding to the three feature points of the image features in the foreground image are selected.

[0059] In step S35, the region where the line connecting the three corresponding feature points in the second feature overlaps with the region where the line connecting the three feature points in the foreground image is the matching feature region.

[0060] In step S36, based on the image features of the foreground image, the matching feature region is fused into the image features of the foreground image to obtain image features about the worker.

[0061] In this invention, during feature fusion, three feature points from the image features in the foreground image can be selected. Based on these three feature points, three feature points from the corresponding second feature can be selected. After obtaining the feature points, the overlapping area between the region connecting the corresponding three feature points in the second feature and the region connecting the three feature points in the foreground image is considered as the matching feature region. After obtaining the matching feature region, it can be fused into the image features of the foreground image, thereby obtaining image features related to the worker. This feature fusion can fuse features that are largely the same in two features, making the overlapping parts more accurate, thus facilitating subsequent action detection.

[0062] On the other hand, the present invention also provides a substation behavior detection system based on video stream analysis, which includes a human face video acquisition module and a detection module. The human face video acquisition module is used to acquire a video stream about the staff. The detection module is used to execute the substation behavior detection method based on video stream analysis as described above based on the acquired video stream.

[0063] On the other hand, the present invention also provides a processor for running a program, wherein the program is executed to perform: the substation behavior detection method based on video stream analysis as described above.

[0064] Through the above technical solution, the substation behavior detection method, system, and processor based on video stream analysis provided by this invention acquires video streams of substation workers' actions, and then performs frame extraction processing on the video streams to obtain images containing workers. After obtaining the images, target detection can be performed on the obtained images to obtain image features related to workers. After obtaining the image features, action analysis and behavior detection can be performed on the acquired image features, thereby generating abnormal alarm information when violations are detected in the images. After receiving the abnormal alarm information, the abnormal warning information about workers and the corresponding video clips can be sent through IoT sensing devices or information systems to achieve alarm push. This substation behavior detection method can accurately identify violations by workers in substations.

[0065] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0066] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0067] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0068] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0069] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0070] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0071] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0072] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0073] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A substation behavior detection method based on video stream analysis, characterized in that, The substation behavior detection method includes: Obtain video streams from the substation showing the actions of staff; The video stream is subjected to frame extraction to obtain images containing staff members; Target detection is performed on the image to obtain image features related to the staff; The image features are subjected to motion analysis and behavior detection to generate abnormal warning information; The system will push alerts by sending notifications of abnormal warnings about staff and corresponding video clips through IoT sensing devices or information systems.

2. The substation behavior detection method according to claim 1, characterized in that, The video stream is subjected to frame extraction to obtain an image containing the staff, including: The video stream is acquired, and frames are extracted from the video stream at a fixed frame extraction frequency; Determine whether the extracted frame image contains staff members; If the image obtained by frame extraction contains workers, increase the frame extraction frequency; If the image obtained by frame extraction does not contain workers, reduce the frame extraction frequency; Collect all the images containing staff obtained from frame extraction.

3. The substation behavior detection method according to claim 1, characterized in that, Perform target detection on the image to obtain image features related to the staff, including: Acquire an image containing staff members, and perform background segmentation on the image to obtain a filtered foreground image containing staff members; Based on the acquired foreground image, edge optimization is performed on the objects in the foreground image to obtain an optimized foreground image; Determine whether any objects are missing in the optimized foreground image; If any parts are missing, repair them. The repaired foreground image and the foreground image without missing parts are subjected to feature transformation to obtain the image features of the foreground image; Obtain the second feature of the image containing the worker from different frames corresponding to the image feature location information in the foreground image; The image features and second features of the foreground image are fused to obtain image features about the staff.

4. The substation behavior detection method according to claim 3, characterized in that, Acquire images containing staff members, and perform background segmentation on the images to obtain filtered foreground images containing staff members, including: Obtain an image containing the staff member, and feed the image into RestNet50 for feature extraction and straighten the extracted features; The straightened features are combined with the positional encoding to obtain the sequence features, which are then input into the Transformer encoder. The Transformer encoder learns global information based on the input sequence features and models the global context through a self-attention mechanism. Then, the Transformer decoder generates prediction boxes. The best match between the predicted bounding box and the true value obtained by formula (1): Formula (1), in, This represents the optimal permutation. Indicates the number of prediction boxes. This represents a set of predicted bounding boxes and actual values. Indicates the label, Represents the actual value. Indicates the prediction box. This represents the cost of matching a single true value with a predicted bounding box. The loss function is calculated to find the best match between the predicted bounding box and the ground truth, and the parameters in the RestNet50, Transformer encoder, and Transformer decoder are adjusted based on minimizing the loss function. The predicted bounding boxes are obtained based on the adjusted parameters. The confidence scores of the predicted bounding boxes are calculated, and the predicted bounding boxes with confidence scores greater than a preset threshold are selected as the foreground images.

5. The substation behavior detection method according to claim 3, characterized in that, Based on the acquired foreground image, edge optimization is performed on the objects in the foreground image to obtain an optimized foreground image, including: Acquire a foreground image containing the staff, and locate the pixel outline boundaries in the foreground image; Based on the obtained pixel outline boundary, select the pixel points of the pixel outline boundary, and with the selected pixel point as the center, obtain all pixel points within the 5*5 neighborhood. Based on the obtained pixels, the initial weight of each pixel is obtained using formula (2): Formula (2), in, Represents pixels The initial weights, The pixel representing the center point, Indicates the weighting parameter; The initial weights of each obtained pixel are normalized, and the pixel value of each pixel is multiplied by the normalized initial weights to update the pixel value of the center pixel. The pixels of all pixel boundaries in the foreground image are traversed to perform edge optimization on the objects in the foreground image, so as to obtain the optimized foreground image.

6. The substation behavior detection method according to claim 3, characterized in that, The foreground image undergoes feature transformation to obtain its image features, including: The foreground image is acquired, and features are extracted from the foreground image to obtain feature points related to the foreground image; Multiple composite geometric transformations are performed on the foreground image to obtain foreground images of various shapes and angles; Features of the foreground image after the composite geometric transformation are extracted to obtain new feature points; The new feature points are mapped back to the feature points of the initial foreground image to obtain aggregated feature points, and the image features of the foreground image are obtained based on the aggregated feature points.

7. The substation behavior detection method according to claim 6, characterized in that, The image features and second features of the foreground image are fused to obtain image features about the worker, including: Select three feature points from the image features in the foreground image; Select three feature points from the second feature corresponding to the three feature points of the image features in the foreground image; The region where the line connecting the three corresponding feature points in the second feature overlaps with the region where the line connecting the three feature points in the foreground image is the matching feature region. Based on the image features of the foreground image, the matching feature region is fused into the image features of the foreground image to obtain image features about the staff.

8. A substation behavior detection system based on video stream analysis, characterized in that, The substation behavior detection system includes: The facial video acquisition module is used to acquire video streams of staff members; The detection module is used to execute the substation behavior detection method based on video stream analysis as described in any one of claims 1-7 based on the acquired video stream.

9. A processor, characterized in that, Used to run a program, wherein the program is run to execute: the substation behavior detection method based on video stream analysis as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Transformer substation personnel behavior recognition method based on monitoring video time sequence action positioning and anomaly detection

    CN111291699A

  • Transformer substation worker abnormal behavior recognition system based on video monitoring

    CN112565675A

  • Object edge optimization method, system and device in video segmentation and storage medium

    CN113902760A

  • Operation risk early warning method and system of transformer substation

    CN117726163A

  • Electric power special operation field intelligent screening method and system

    CN119580160A