Substation binocular vision hidden danger behavior warning method, system, device and medium
By using binocular cameras and a 3D sparse point cloud ranging algorithm, the distance between potential hazards in substations and equipment can be accurately measured, solving the problem that existing technologies cannot accurately control the hazards of potential hazards and realizing intelligent safety monitoring of substations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAYAN INTELLIGENT TECH (GRP) CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies cannot accurately measure the distance between potential hazards in substations and power equipment, making it impossible to precisely control the degree of harm caused by these hazards to the equipment.
Binocular cameras are used to acquire binocular images of key nodes of the equipment, and the spatial relationship of the equipment in the coordinate system of the left eye camera is constructed. Combined with personnel behavior detection and segmentation models, the distance between personnel and equipment is calculated by a three-dimensional sparse point cloud ranging algorithm, and the hazard alarm level is determined based on the distance.
It enables differentiated alarms for potential hazards in substations, improving the intelligence level and response efficiency of operational safety monitoring.
Smart Images

Figure CN122135505A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of behavioral alarm technology, and in particular to a method, system, device and medium for binocular visual behavioral alarm of potential hazards in substations. Background Technology
[0002] During power line inspections, construction work, and equipment maintenance, substation personnel often engage in hazardous behaviors such as not wearing safety helmets or work clothes, making phone calls, or smoking, which can cause abnormal operation of substation equipment and endanger the lives of substation personnel. Therefore, effectively monitoring hazardous behaviors in substations, accurately measuring the distance between hazardous behaviors and substation equipment, and promptly eliminating the hazards posed by these behaviors have become urgent expectations and goals for substation hazardous behavior monitoring.
[0003] Existing methods for monitoring potential hazards in substations, such as image-based target detection algorithms like the YOLO and DETR series, can only monitor potential hazards in substations and cannot directly measure the distance between the hazard and the substation equipment, thus failing to accurately control the degree of harm caused by the hazard. Similarly, video-based target tracking algorithms, such as inter-frame difference algorithms and optical flow tracking algorithms, can also only monitor potential hazards in substations and cannot directly measure the distance between the hazard and the substation equipment, thus failing to accurately control the degree of harm caused by the hazard. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to overcome the shortcomings of the prior art and provide a method, system, device and medium for detecting potential hazards in substations using binocular vision.
[0005] This invention provides the following technical solution: In a first aspect, the present invention provides a method for detecting potential safety hazards in a substation using binocular vision, the method comprising: The binocular images of multiple key equipment nodes in a substation are acquired using a binocular camera, and the spatial relationship of each key equipment node in the coordinate system of the left camera is constructed based on the binocular images. The binocular camera is used to sample the current substation binocular visual image at intervals. The current substation binocular visual image is input into the personnel behavior detection model to obtain the current binocular behavior detection result. If there are hidden danger behavior rectangles in the current binocular behavior detection result, the binocular camera is used to collect the real-time substation binocular visual image. The real-time binocular visual image of the substation is input into the personnel behavior detection model to obtain the real-time binocular behavior detection result. If the real-time binocular behavior detection result contains the hazard behavior rectangle, the real-time binocular visual image of the substation is input into the personnel entity segmentation model to obtain the left eye human body segmentation result and the right eye human body segmentation result. Based on the human body segmentation results of the left and right eyes, human body key nodes are matched to obtain the three-dimensional coordinates of multiple human body key nodes. Based on the three-dimensional coordinates of each human body key node, the spatial relationship of each human body key node in the coordinate system of the left eye camera is constructed. Based on the spatial relationship of each key human node in the coordinate system of the left eye camera and the spatial relationship of each key device node in the coordinate system of the left eye camera, three-dimensional sparse point cloud ranging is performed to obtain the distance between the person and the device. The hazard alarm level is determined based on the distance between personnel and equipment and the preset distance index, and the hazard behavior of the substation is given differentiated alarms based on the hazard alarm level.
[0006] Secondly, the present invention provides a binocular vision-based hazard behavior alarm system for substations, the system comprising: The equipment space construction module is used to acquire binocular images of multiple key equipment nodes in a substation through a binocular camera, and construct the spatial relationship of each key equipment node in the coordinate system of the left eye camera based on the binocular images. The hazard behavior detection module is used to periodically sample the current substation binocular visual images using the binocular camera, input the current substation binocular visual images into the personnel behavior detection model, and obtain the current binocular behavior detection results. If there are hazard behavior rectangles in the current binocular behavior detection results, then the binocular camera is used to collect real-time substation binocular visual images; the real-time substation binocular visual images are input into the personnel behavior detection model to obtain real-time binocular behavior detection results. The personnel entity segmentation module is used to input the real-time binocular visual image of the substation into the personnel entity segmentation model if the real-time binocular behavior detection results all contain the rectangle of the hidden danger behavior, so as to obtain the left-eye human body segmentation result and the right-eye human body segmentation result. The human body space construction module is used to match key human body nodes based on the left-eye human body segmentation results and the right-eye human body segmentation results, obtain the three-dimensional coordinates of multiple key human body nodes, and construct the spatial relationship of each key human body node in the left-eye camera coordinate system based on the three-dimensional coordinates of each key human body node. The three-dimensional point cloud ranging module is used to perform three-dimensional sparse point cloud ranging based on the spatial relationship of each of the key human body nodes in the left eye camera coordinate system and the spatial relationship of each of the key equipment nodes in the left eye camera coordinate system, so as to obtain the distance between the personnel and the equipment. The hazard differentiation alarm module is used to determine the hazard alarm level based on the distance between personnel and equipment and a preset distance indicator, and to issue differentiated alarms for hazard behaviors in the substation based on the hazard alarm level.
[0007] Thirdly, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the substation binocular visual hazard behavior alarm method as described in any of the foregoing embodiments.
[0008] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the substation binocular visual hazard behavior alarm method as described in any of the foregoing embodiments.
[0009] This invention discloses a method, system, device, and medium for substation binocular visual hazard behavior alarm. It acquires binocular images of key equipment nodes using a binocular camera, thereby constructing the spatial relationships between these nodes. By periodically sampling the substation's binocular visual images, it detects binocular hazard behavior bounding boxes, and then acquires real-time substation binocular visual images. Using these real-time images, it detects the bounding boxes and constructs the spatial relationships of key human body nodes using binocular human body segmentation and key point matching. Based on the spatial relationships of these nodes and the equipment's key nodes, it uses a sparse point cloud ranging algorithm to calculate the distance between personnel and equipment, and then provides differentiated alarms for different hazard levels. By establishing the spatial relationship between key human body nodes and key equipment nodes in a unified left-eye camera coordinate system, and quantifying the distance between hazard behavior and substation equipment, it achieves differentiated alarms for substation hazard behavior, effectively improving the intelligence level and response efficiency of operational safety monitoring. Attached Figure Description
[0010] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope of protection of the present invention. In the various drawings, similar components are numbered similarly.
[0011] Figure 1 This embodiment shows a flowchart of the substation binocular vision-based hazard behavior alarm method. Figure 2This embodiment shows a schematic diagram illustrating the spatial relationship of key nodes of the construction device in the left eye camera coordinate system. Figure 3 This diagram illustrates the spatial relationship of the key nodes of the device proposed in this embodiment in the coordinate system of the left eye camera. Figure 4 This diagram illustrates the network structure of the personnel behavior detection model proposed in this embodiment. Figure 5 This embodiment illustrates a flowchart of constructing the spatial relationships of key human body nodes in the left eye camera coordinate system. Figure 6 This embodiment illustrates a flowchart of the three-dimensional sparse point cloud ranging method proposed in this example. Figure 7 This embodiment illustrates a three-dimensional spatial relationship between the human body and the device. Figure 8 A schematic diagram of the substation binocular vision hazard behavior alarm system proposed in this embodiment is shown.
[0012] Explanation of reference numerals in the attached diagram: 800-Substation binocular vision hazard behavior alarm system; 801-Equipment spatial structure module; 802-Hazard behavior detection module; 803-Personnel entity segmentation module; 804-Human body spatial structure module; 805-3D point cloud ranging module; 806-Hazard difference alarm module. Detailed Implementation
[0013] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0014] The components of the embodiments of the invention described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0015] In the following, the terms “comprising,” “having,” and their cognates, which may be used in various embodiments of the invention, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as excluding, firstly, the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more features, numbers, steps, operations, elements, components, or combinations thereof.
[0016] Furthermore, the terms "first," "second," and "third" are used only to distinguish descriptions and should not be interpreted as indicating or implying relative importance.
[0017] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of the invention pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be interpreted as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of the invention.
[0018] Example 1 This disclosure provides a method for detecting potential safety hazards in substations using binocular vision.
[0019] Please see Figure 1 The binocular visual hazard behavior alarm method for this substation includes steps S101 to S106, and each step is described in detail below.
[0020] Step S101: Obtain binocular images of multiple key equipment nodes in the substation using a binocular camera, and construct the spatial relationship of each key equipment node in the coordinate system of the left eye camera based on the binocular images.
[0021] In this embodiment, binocular images of key equipment nodes in a substation are acquired by using a binocular camera to obtain binocular images of the equipment.
[0022] Furthermore, based on the pixel coordinates of key points of the device in the binocular image, the three-dimensional coordinates are calculated, thereby using the three-dimensional coordinate information to construct the spatial relationship of key nodes of the device in the coordinate system of the left eye camera, thus providing a fixed reference and judgment basis for subsequent spatial comparison between the human body position and the device position.
[0023] It should be noted that before acquiring images using a binocular camera, the binocular camera is calibrated and corrected to correct two non-coplanar aligned images into two coplanar aligned images, thereby obtaining a binocular camera with the left eye camera coordinate system as the reference.
[0024] Understandably, calibration and correction eliminate the distortion of the camera itself and determine the relative positional relationship between the left and right cameras. This provides a precise image data source for subsequent processing, ensuring that the image data collected from the source is accurate.
[0025] In one specific embodiment, please refer to Figure 2 The diagram shows a schematic representation of the spatial relationship of key nodes of the construction device in the coordinate system of the left eye camera. The device's binocular image includes the device's left eye image and the device's right eye image. Step S101 includes steps S1011 to S1014. Each step is described in detail below.
[0026] Step S1011: For each of the device key nodes, calculate the horizontal coordinate difference based on the pixel coordinates of the device key node in the left eye image and the pixel coordinates of the device key node in the right eye image to obtain the parallax of the device key node.
[0027] In this embodiment, the vertical coordinates of key nodes in the left and right view images of the device are the same, while the horizontal coordinates create parallax. Therefore, the pixel coordinates of the key nodes in the left view image of the device ( , ), and the device key node pixel coordinates of the right eye image of the device ( , The parallax of critical equipment nodes is obtained by calculating the difference in the horizontal coordinates. For example, the parallax of critical equipment nodes. .
[0028] Step S1012: Based on the intrinsic parameter matrix and baseline length of the binocular camera, the disparity of the key nodes of the device, and the pixel coordinates of the key nodes of the device in the left eye image, calculate the three-dimensional coordinates of the key nodes of the device in the coordinate system of the left eye camera to obtain the three-dimensional coordinates of the key nodes of the device.
[0029] In this embodiment, the intrinsic parameter matrix of the stereo camera is used. and baseline length B And the parallax of the device's key nodes will determine the pixel coordinates of the device's key nodes in the left eye image. , The coordinates of the key nodes of the device are converted into three-dimensional coordinates in the coordinate system of the left eye camera to obtain the three-dimensional coordinates of the key nodes of the device.
[0030] As an example, the intrinsic parameter matrix of a stereo camera after calibration and correction. In the formula, , This represents the camera pixel focal length, which is a scaling factor that converts three-dimensional distance in the camera coordinate system into the number of image pixels. Since pixels are generally square, therefore... ;( , The principal point coordinates are the coordinates of the intersection of the camera's optical axis and the imaging plane. This intersection point is usually located at the center of the image in the image pixel coordinate system.
[0031] Meanwhile, the extrinsic translation vector of the right camera relative to the left camera is And baseline length B Let be the magnitude of the translation vector, i.e.: .
[0032] Three-dimensional coordinates of key equipment nodes , , The conversion formula for ) is: .
[0033] Step S1013: Calculate the spatial plane equation and spatial line segment equation based on the three-dimensional coordinates of the key nodes of the equipment to obtain the spatial plane and spatial line segment corresponding to the key nodes of the equipment.
[0034] In this embodiment, the spatial plane equation and spatial line segment equation are calculated using the three-dimensional coordinates of each key node of the equipment to obtain the spatial plane and spatial line segment associated with the key node of the corresponding equipment.
[0035] Step S1014: Bind the key nodes of the device according to their three-dimensional coordinates, the spatial planes and spatial line segments corresponding to the key nodes of the device, and obtain the spatial relationship of the key nodes of the device in the coordinate system of the left eye camera.
[0036] In this embodiment, each key node of the device and its associated spatial plane and spatial line segment are bound together, and each key node of the device and its bound spatial plane and spatial line segment are stored and recorded. The spatial relationship of the key nodes of the device in the left eye camera coordinate system is constructed using the left eye camera spatial coordinate system.
[0037] Please see Figure 3 The diagram shows the spatial relationship of key device nodes in the left-eye camera coordinate system. The pixel coordinates of the key device nodes in the left-eye device image (their coordinates are a1~a6, and the number of key nodes varies for different devices) are obtained after using the key node 3D coordinate transformation formula.
[0038] Next, the spatial plane equations are calculated using the three-dimensional coordinates (A1~A6) of the key equipment nodes to obtain the spatial planes of the key equipment nodes (such as planes A1A2A3A4 and A3A4A5A6). At the same time, the spatial line segment equations are calculated using the three-dimensional coordinates (A1~A6) of the key equipment nodes to obtain the spatial line segments of the key equipment nodes (such as line segments A1A2, A2A3, A3A4, A1A4, A4A5, A5A6, and A3A6). Next, the key nodes of the equipment are bound to the corresponding spatial planes and spatial line segments (spatial node A1 is bound to planes A1A2A3A4, line segment A1A2, and line segment A1A4; spatial node A2 is bound to planes A1A2A3A4, line segment A1A2, and line segment A2A3; spatial node A3 is bound to planes A1A2A3A4, planes A3A4A5A6, line segment A2A3, line segment A3A4, and line segment A3A6; spatial node A4 is bound to planes A1A2A3A4, planes A3A4A5A6, line segment A1A4, line segment A3A4, and line segment A4A5; spatial node A5 is bound to planes A3A4A5A6, line segment A4A5, and line segment A5A6; spatial node A6 is bound to planes A3A4A5A6, line segment A3A6, and line segment A5A6). Finally, the three-dimensional coordinates of each key node of the device, as well as the bound spatial planes and spatial line segments, are stored and recorded.
[0039] Step S102: Use the binocular camera to sample the current substation binocular visual image at intervals, input the current substation binocular visual image into the personnel behavior detection model, and obtain the current binocular behavior detection result. If there are hidden danger behavior rectangles in the current binocular behavior detection result, then use the binocular camera to collect the real-time substation binocular visual image.
[0040] In this embodiment, a binocular camera is used to sample the current binocular visual image of the substation at a preset time interval. The current binocular visual image of the substation is then input into the personnel behavior detection model for preliminary hazard behavior detection to obtain the current binocular behavior detection result. If there is a hazard behavior rectangle in the current binocular behavior detection result, it is considered that a binocular hazard behavior has occurred at the current moment. The binocular camera is then used to collect real-time binocular visual images of the substation for detailed analysis.
[0041] In one specific embodiment, the current binocular behavior detection result includes a left-eye behavior detection result and a right-eye behavior detection result. Step S102 includes: inputting the left-eye image of the substation in the current binocular visual image of the substation into the personnel behavior detection model to obtain the left-eye behavior detection result; if there is no hazard behavior rectangle in the left-eye behavior detection result, then re-execute the step of sampling the current substation binocular visual image at intervals using the binocular camera; if there is a hazard behavior rectangle in the left-eye behavior detection result, then inputting the right-eye image of the substation in the current binocular visual image of the substation into the personnel behavior detection model to obtain the right-eye behavior detection result; if there is a hazard behavior rectangle in the right-eye behavior detection result, then using the binocular camera to collect the real-time substation binocular visual image; if there is no hazard behavior rectangle in the right-eye behavior detection result, then re-execute the step of sampling the current substation binocular visual image at intervals using the binocular camera.
[0042] In this embodiment, the left eye image of the substation in the current binocular vision image of the substation is input into the personnel behavior detection model to obtain the left eye behavior detection result (including multiple behavior rectangles and corresponding behavior types).
[0043] If no potential danger behavior rectangle is found in the left eye behavior detection result, it is assumed that no potential danger behavior has occurred at the current moment, and the next pair of interval sampling binocular visual images of the substation are obtained.
[0044] If the left-eye behavior detection result contains a bounding box of a potential danger, it is considered that a potential danger has occurred at the current moment. The right-eye image is then input into the personnel behavior detection model to obtain the right-eye behavior detection result (which includes multiple bounding boxes of behaviors and their corresponding behavior types).
[0045] If there is no hidden danger behavior bounding box in the right eye behavior detection result, it is considered that no binocular hidden danger behavior has occurred at the current moment (at this time, the target may be located in the image boundary area or the model may generate false alarm information), and the next pair of interval sampling binocular visual images of the substation are obtained.
[0046] If a hazard behavior rectangle is found in the right eye behavior detection result, it is considered that a binocular hazard behavior is occurring at the current moment, and real-time binocular visual images of the substation are captured.
[0047] Step S103: Input the real-time binocular visual image of the substation into the personnel behavior detection model to obtain the real-time binocular behavior detection result. If the real-time binocular behavior detection result contains the hazard behavior rectangle, then input the real-time binocular visual image of the substation into the personnel entity segmentation model to obtain the left-eye human body segmentation result and the right-eye human body segmentation result.
[0048] In this embodiment, real-time binocular visual images of the substation are input into a personnel behavior detection model to obtain real-time binocular behavior detection results. If both the left and right eye behavior detection results contain bounding boxes of potential hazards, it is determined that a potential hazard exists at the substation at the current moment. The real-time binocular visual images of the substation are then input into a personnel entity segmentation model to obtain left and right eye human body segmentation results. Both left and right eye human body segmentation results include multiple human body masks.
[0049] Understandably, by using a personnel behavior detection model for initial screening, fine-grained analysis is only triggered when suspected potential hazards are detected consecutively. This effectively reduces the overall computational load of the system, improves resource utilization, and avoids the waste of computing power caused by real-time complex segmentation calculations. Simultaneously, the instance segmentation technique can accurately distinguish the outline of the human body at the pixel level, unaffected by changes in lighting, shadows, or complex backgrounds. This provides a more refined and robust input for subsequent extraction of key human body nodes than target detection boxes, significantly improving the accuracy of key human body node localization.
[0050] In one specific embodiment, please refer to Figure 4 The diagram shows the network structure of the personnel behavior detection model, which is an improved YOLO12-m object detection model. The improved YOLO12-m object detection model includes a backbone network, a feature pyramid network (PA-SLAM-ECAM-FPN), and a head detection network. The feature pyramid network includes a first-scale interactive attention module (SIAM1), a second-scale interactive attention module (SIAM2), a first convolutional layer (Conv1), a second convolutional layer (Conv2), a first efficient convolutional attention module (ECAM1), a second efficient convolutional attention module (ECAM2), and a third efficient convolutional attention module (ECAM3). The head detection network includes a first classification regression module (regression1 and classification1), a second classification regression module (regression2 and classification2), and a third classification regression module (regression3 and classification3).
[0051] Specifically, the backbone network includes Conv3, Conv4, and Conv5. Taking the left eye image of the substation in the current binocular vision image of the substation as an example, the left eye image of the substation is used as the input of the backbone network to obtain the first layer feature map F3 corresponding to the output of Conv3, the second layer feature map F4 corresponding to the output of Conv4, and the third layer feature map F5 corresponding to the output of Conv5.
[0052] Furthermore, the third-layer feature map F5 and the second-layer feature map F4 are input into the first-scale interactive attention module to obtain the first-scale interactive attention feature map SIAF4. The first-scale interactive attention feature map SIAF4 is then concatenated with the second-layer feature map F4 to obtain the first fusion map ~P4.
[0053] Furthermore, the first fused image ~P4 and the first layer feature map F3 are input into the second scale interactive attention module to obtain the second scale interactive attention feature map SIAF3, and the second scale interactive attention feature map SIAF3 is concatenated with the first layer feature map F3 to obtain the second fused image P3.
[0054] Furthermore, the second fused image P3 is input into the first convolutional layer to obtain the first feature map PAF4 with the same resolution as ~P4, and the first feature map PAF4 is concatenated with the first fused image ~P4 to obtain the third fused image P4.
[0055] Furthermore, the third fusion map P4 is input into the second convolutional layer to obtain the second feature map PAF5 with the same resolution as F5. The second feature map PAF5 is then concatenated with the third layer feature map F5 to obtain the fourth fusion map P5.
[0056] Furthermore, the second fused image P3, the third fused image P4, and the fourth fused image P5 are subjected to efficient convolutional attention feature extraction through the first efficient convolutional attention module, the second efficient convolutional attention module, and the third efficient convolutional attention module, respectively, to obtain cross-level features from top to bottom and bottom to top, including the first efficient convolutional attention feature map T3, the second efficient convolutional attention feature map T4, and the third efficient convolutional attention feature map T5.
[0057] Furthermore, the first efficient convolutional attention feature map T3, the second efficient convolutional attention feature map T4, and the third efficient convolutional attention feature map T5 are input into the first classification regression module, the second classification regression module, and the third classification regression module, respectively, to obtain the first initial monocular behavior detection result, the second initial monocular behavior detection result, and the third initial monocular behavior detection result.
[0058] Furthermore, the left-eye behavior detection result in the current binocular behavior detection result is obtained based on the first initial monocular behavior detection result, the second initial monocular behavior detection result, and the third initial monocular behavior detection result.
[0059] It should be noted that the right eye behavior detection results in the current binocular behavior detection results, as well as the left eye behavior detection results and right eye behavior detection results in the real-time binocular behavior detection results, are all referenced separately. Figure 4 The network structure was detected.
[0060] It should be further explained that the process involves obtaining the original image library of substation personnel behavior, constructing an augmented image library of substation personnel behavior (augmentation methods include left and right flipping, angle rotation, Gaussian noise, image blending, etc.), merging the original image library of substation personnel behavior and the augmented image library of substation personnel behavior to obtain the training sample library of substation personnel behavior, and training the improved YOLO12-m target detection algorithm based on the training sample library of substation personnel behavior to obtain the personnel behavior detection model.
[0061] Simultaneously, the original image library of substation personnel entities is obtained, and an augmented image library of substation personnel entities is constructed (the augmentation methods include left and right flipping, angle rotation, Gaussian noise, etc.); the original image library of substation personnel entities and the augmented image library of substation personnel entities are merged to obtain a training sample library of substation personnel entities; based on the training sample library of substation personnel entities, the YOLO12-m instance segmentation algorithm is trained to obtain a personnel entity segmentation model.
[0062] Step S104: Perform human body key node matching based on the left-eye human body segmentation result and the right-eye human body segmentation result to obtain the three-dimensional coordinates of multiple human body key nodes, and construct the spatial relationship of each human body key node in the left-eye camera coordinate system based on the three-dimensional coordinates of each human body key node.
[0063] In this embodiment, the ORB algorithm is used to match key human body nodes in the left and right human body segmentation results to obtain the three-dimensional coordinates of multiple key human body nodes. Furthermore, the spatial relationship of each key human body node in the left camera coordinate system is constructed based on the three-dimensional coordinates of each key human body node.
[0064] In one specific embodiment, please refer to Figure 5 The diagram shows a process for constructing the spatial relationship of key human body nodes in the left eye camera coordinate system. Step S104 includes steps S1041 to S1044, and each step is described in detail below.
[0065] Step S1041: Perform mask matching based on the left-eye human body segmentation result and the right-eye human body segmentation result to obtain the binocular human body key node pixel coordinates of each of the human body key nodes.
[0066] In this embodiment, the ORB algorithm is used to perform human body mask matching based on the left and right human body segmentation results to obtain the binocular key nodes that match the left and right human body masks, thereby obtaining the binocular human body key node pixel coordinates of each human body key node.
[0067] Step S1042: For each of the human key nodes, calculate the difference in horizontal coordinates based on the left and right eye pixel coordinates of the human key node in the binocular human key node pixel coordinates to obtain the parallax of the human key node.
[0068] In this embodiment, for each key human body node, the horizontal coordinate difference is calculated based on the left and right eye key human body node pixel coordinates in the binocular key human body node pixel coordinates to obtain the parallax of the key human body node.
[0069] Step S1043: Based on the intrinsic parameter matrix and baseline length of the binocular camera, as well as the disparity of the human body key node and the pixel coordinates of the left human body key node, calculate the three-dimensional coordinates of the human body key node in the coordinate system of the left camera to obtain the three-dimensional coordinates of the human body key node.
[0070] In this embodiment, the pixel coordinates of the human key node in the left eye are transformed using the intrinsic parameter matrix and baseline length of the binocular camera, as well as the parallax of the human key node, to obtain the three-dimensional coordinates of the human key node in the coordinate system of the left eye camera, i.e., the three-dimensional coordinates of the human key node. The transformation formula for the three-dimensional coordinates of the human key node is the same as the transformation formula for the three-dimensional coordinates of the device key node.
[0071] Step S1044: Project the three-dimensional coordinates of each of the key human body nodes onto the left eye camera coordinate system to obtain the spatial relationship of each of the key human body nodes in the left eye camera coordinate system.
[0072] In this embodiment, the three-dimensional coordinates of each key human body node are projected onto the left eye camera coordinate system to obtain the spatial relationship of each key human body node on the left eye camera coordinate system, thereby realizing the left eye camera coordinate system as a unified measurement benchmark and providing a basis for subsequent spatial alignment.
[0073] Step S105: Based on the spatial relationship of each key human node in the left eye camera coordinate system and the spatial relationship of each key device node in the left eye camera coordinate system, perform three-dimensional sparse point cloud ranging to obtain the distance between the person and the device.
[0074] In this embodiment, by combining the spatial relationships of each key human body node in the left eye camera coordinate system and the spatial relationships of each key equipment node in the left eye camera coordinate system, three-dimensional sparse point cloud ranging is performed. By calculating the distance between the key human body node and the key equipment node, the distance between the personnel and the equipment is obtained, thereby achieving precise quantification of the distance between the potential danger behavior and the power equipment.
[0075] In one specific embodiment, please refer to Figure 6The diagram shown illustrates a process for ranging a three-dimensional sparse point cloud. Step S105 includes steps S1051 to S1058, and each step is described in detail below.
[0076] Step S1051: Based on the spatial relationships of each of the key human body nodes in the left eye camera coordinate system and the spatial relationships of each of the key device nodes in the left eye camera coordinate system, obtain the three-dimensional spatial relationship between the human body and the device.
[0077] In this embodiment, the spatial relationships of key human body nodes in the left eye camera coordinate system and the spatial relationships of key device nodes in the left eye camera coordinate system are integrated to obtain the three-dimensional spatial relationship between the human body and the device. This makes it more intuitive to calculate the spatial topological relationship between the two and provides a data structure foundation for quantitative analysis.
[0078] Step S1052: For each of the key human body nodes, based on the three-dimensional spatial relationship between the human body and the device, calculate the distance from each key human body node to all the key device nodes to obtain multiple node distances.
[0079] In this embodiment, for each key human body node, based on the three-dimensional coordinates in the three-dimensional spatial relationship between the human body and the device, the distance from the key human body node to each key device node is calculated, thus obtaining the distances between multiple nodes corresponding to the key human body node.
[0080] Understandably, by traversing all combinations of key human body nodes and key equipment nodes, a complete distance candidate set is constructed to ensure that subsequent analysis does not miss any key point pairs that may lead to dangerous contact.
[0081] Step S1053: The minimum value among the multiple node distances is taken as the candidate distance, and the human body key node and the device key node corresponding to the candidate distance are marked as human body candidate node and device candidate node, respectively.
[0082] In this embodiment, the minimum value among the distances of each node corresponding to the human body key node is taken as the candidate distance corresponding to the human body key node, which means the most dangerous distance corresponding to the corresponding human body key point. The human body key node and the device key node corresponding to the candidate distance are marked as human body candidate node and device candidate node, respectively.
[0083] Step S1054: Based on the three-dimensional spatial relationship between the human body and the device, calculate the projection node coordinates from the human body candidate node to the spatial plane bound to the device candidate node.
[0084] In this embodiment, based on the three-dimensional coordinates and spatial plane in the three-dimensional spatial relationship between the human body and the device, the projection node coordinates of the human body candidate node to the spatial plane bound to the device candidate node are calculated. The spatial plane bound to the device candidate node more realistically reflects the physical surface of the device, providing an accurate geometric reference for calculating the actual distance from the human body to the device surface, and avoiding the error caused by simplifying the device to a point.
[0085] Step S1055: If the coordinates of the projected node are within the area of the spatial plane bound to the device candidate node, then based on the three-dimensional spatial relationship between the human body and the device, calculate the distance from the human body candidate node to the spatial plane bound to the device candidate node, and use it as the distance of the human body key node.
[0086] In this embodiment, based on the three-dimensional spatial relationship between the human body and the device, it is determined whether the coordinates of the projection node are within the area of the spatial plane bound to the candidate node of the device.
[0087] If the coordinates of the projected node are within the area of the spatial plane bound to the device candidate node, then it can be known that the human body key node is directly facing the spatial plane bound to the device key node. The vertical distance from the point to the plane is directly calculated, reflecting the real spatial relationship between the human body and the device surface. Therefore, based on the three-dimensional spatial relationship between the human body and the device, the vertical distance from the human body candidate node to the spatial plane bound to the device candidate node is calculated as the distance between the human body key node and the device, i.e., the human body key node distance.
[0088] Step S1056: If the coordinates of the projection node are not within the area of the spatial plane bound to the device candidate node, then based on the three-dimensional spatial relationship between the human body and the device, calculate the distance from the human body candidate node to the spatial line segment bound to the device candidate node to obtain the point-line distance.
[0089] In this embodiment, if the coordinates of the projected node are not within the area of the spatial plane bound to the device candidate node, it can be known that the key human body node is close to the edge or corner of the device. Directly calculating the distance from the human point to the line segment can accurately describe the contact risk between the human body and the edge of the device, making up for the insufficiency of planar distances in covering edge situations, and ensuring the comprehensiveness and accuracy of distance calculations in various spatial locations. Therefore, based on the three-dimensional spatial relationship between the human body and the device, the perpendicular distance from the human body candidate node to the spatial line segment bound to the device candidate node is calculated to obtain the point-to-line distance.
[0090] Step S1057: The minimum value between the point-line distance and the candidate distance is taken as the distance of the human body key node corresponding to the human body key node.
[0091] In this embodiment, the minimum value between the point-line distance and the candidate distance is taken as the distance between the human body key node and the device, i.e., the human body key node distance.
[0092] Step S1058: Discard the long and short distances from each of the key human body node distances according to a preset ratio to obtain the remaining key human body node distances. Calculate the average of the remaining key human body node distances to obtain the distance between the person and the equipment.
[0093] In this embodiment, the long and short distances of each key human body node are discarded according to a preset ratio to obtain the remaining key human body node distances. The average of the remaining key human body node distances is calculated to obtain the distance between the personnel and the equipment, thereby effectively ensuring the accuracy of the distance between the personnel and the equipment.
[0094] For example, please see Figure 7 The diagram shows the three-dimensional spatial relationship between the human body and the equipment. The two-dimensional image coordinates a1~a6 of the key nodes of the equipment are transformed into three-dimensional coordinates A1~A6, and the two-dimensional image coordinates p1~p5 of the key nodes of the human body are transformed into three-dimensional coordinates P1~P5.
[0095] Based on the positional relationship between the key human body nodes P1~P5 and the key equipment nodes A1~A6, first calculate the distance from each key human body node to all key equipment nodes (for example, the distance from P2 to all key equipment nodes P2A1, P2A2, P2A3, P2A4, P2A5, P2A6).
[0096] For each human body key node, determine the human body key node and the device key node with the minimum distance (e.g., P2, A3), and mark the human body key node and the device key node with the minimum distance as human body candidate node (e.g., P2) and device candidate node (e.g., A3) respectively, and mark the minimum distance as candidate distance (e.g., distance P2A3).
[0097] Next, calculate the projection node coordinates (e.g., Q2) from the human candidate node (e.g., P2) to the spatial plane (e.g., plane A3A4A5A6) to which the device candidate node is bound. If the projection node is within the area of the plane to which the device candidate node is bound, then calculate the distance from the human candidate node to the plane to which the device candidate node is bound (e.g., distance P2Q2). This distance is the distance of the human key node.
[0098] If the projection node is not within the area of the plane to which the device candidate node is bound (for example, the projection point of P2 on plane A1A2A3A4 is not within the area of plane A1A2A3A4, but in the extended area of plane A1A2A3A4), then calculate the distance from the human candidate node to the line segment bound to the device candidate node (such as line segment A3A2, line segment A3A4, line segment A3A6), compare the size of this distance with the candidate distance, and select the smallest distance as the distance of the human key node.
[0099] By calculating the distance from each key node of the human body to the device, we obtain the distances of multiple key nodes of the human body, namely: the distance from P1 to the device P1Q1, the distance from P2 to the device P2Q2, the distance from P3 to the device P3Q3, the distance from P4 to the device P4Q4, and the distance from P5 to the device P5Q5.
[0100] Finally, based on multiple human critical node distances (e.g., 5 human critical node distances), according to the proportional relationship, a portion of the long-distance human critical node distances are discarded (e.g., if 20% of the human critical node distances are discarded, then 1 is discarded, such as P5Q5), and a portion of the short-distance human critical node distances are discarded (e.g., if 20% of the human critical node distances are discarded, then 1 is discarded, such as P1Q1). Then, the average value of the remaining human critical node distances (e.g., 60% of the human critical node distances, which is 3, such as P2Q2, P3Q3, and P4Q4) is calculated to obtain the final distance between personnel and equipment.
[0101] Step S106: Determine the hazard alarm level based on the distance between personnel and equipment and the preset distance index, and issue differentiated alarms for the hazard behavior of the substation based on the hazard alarm level.
[0102] In this embodiment, the distance between personnel and equipment is compared with a preset distance index to determine the hazard alarm level corresponding to the distance between personnel and equipment, and the current hazard behavior of the substation is given a differentiated alarm based on the hazard alarm level.
[0103] In one specific embodiment, the preset distance indicators include a first danger distance, a second danger distance, a third danger distance, and a fourth danger distance.
[0104] Specifically, if the distance between personnel and equipment is less than the first danger distance, the hazard alarm level is a first-level alarm; if the distance between personnel and equipment is less than the second danger distance, the hazard alarm level is a second-level alarm; if the distance between personnel and equipment is less than the third danger distance, the hazard alarm level is a first-level warning; if the distance between personnel and equipment is less than the fourth danger distance, the hazard alarm level is a second-level warning; if the distance between personnel and equipment is greater than or equal to the fourth danger distance, no warning is issued.
[0105] The substation binocular vision-based hazard behavior alarm method proposed in this embodiment acquires binocular images of key equipment nodes using a binocular camera, thereby constructing the spatial relationships between these key nodes. It then detects binocular hazard behavior bounding boxes by periodically sampling the substation's binocular vision images, and acquires real-time substation binocular vision images. Using these real-time images, it detects the bounding boxes and constructs the spatial relationships between key human body nodes using binocular human body segmentation and key point matching. Based on the spatial relationships between these key human body nodes and key equipment nodes, a sparse point cloud ranging algorithm is used to calculate the distance between personnel and equipment, and differentiated alarms are issued for different hazard alarm levels accordingly. By establishing the spatial relationship between key human body nodes and key equipment nodes in a unified left-eye camera coordinate system, and quantifying the distance between hazard behavior and substation equipment, differentiated alarms for substation hazard behavior are achieved, effectively improving the intelligence level and response efficiency of operational safety monitoring.
[0106] Example 2 Furthermore, this disclosure provides a substation binocular vision-based hazard behavior alarm system 800; please refer to [link to relevant documentation]. Figure 8 The system includes: The equipment space construction module 801 is used to acquire binocular images of multiple key equipment nodes in a substation through a binocular camera, and construct the spatial relationship of each key equipment node in the left eye camera coordinate system based on the binocular images. The hazard behavior detection module 802 is used to periodically sample the current substation binocular visual image using the binocular camera, input the current substation binocular visual image into the personnel behavior detection model to obtain the current binocular behavior detection result. If there are hazard behavior rectangles in the current binocular behavior detection result, then the binocular camera is used to collect the real-time substation binocular visual image; the real-time substation binocular visual image is input into the personnel behavior detection model to obtain the real-time binocular behavior detection result. The personnel entity segmentation module 803 is used to input the real-time substation binocular visual image into the personnel entity segmentation model if the real-time binocular behavior detection results all contain the hazard behavior rectangle, so as to obtain the left-eye human body segmentation result and the right-eye human body segmentation result. The human body space construction module 804 is used to match key human body nodes based on the left-eye human body segmentation result and the right-eye human body segmentation result, obtain the three-dimensional coordinates of multiple key human body nodes, and construct the spatial relationship of each key human body node in the left-eye camera coordinate system based on the three-dimensional coordinates of each key human body node. The three-dimensional point cloud ranging module 805 is used to perform three-dimensional sparse point cloud ranging based on the spatial relationship of each of the key human body nodes in the coordinate system of the left eye camera and the spatial relationship of each of the key equipment nodes in the coordinate system of the left eye camera, so as to obtain the distance between the personnel and the equipment. The hazard differential alarm module 806 is used to determine the hazard alarm level based on the distance between personnel and equipment and a preset distance indicator, and to issue differentiated alarms for the hazard behavior of the substation based on the hazard alarm level.
[0107] In an optional implementation, the 3D point cloud ranging module 805 is further configured to: obtain the 3D spatial relationship between the human body and the device based on the spatial relationship of each of the human body key nodes in the left eye camera coordinate system and the spatial relationship of each of the device key nodes in the left eye camera coordinate system; calculate the distance from each human body key node to all the device key nodes based on the 3D spatial relationship between the human body and the device, thereby obtaining multiple node distances; take the minimum value among the multiple node distances as a candidate distance, and mark the human body key node and device key node corresponding to the candidate distance as human body candidate node and device candidate node, respectively; calculate the projection node coordinates from the human body candidate node to the spatial plane bound to the device candidate node based on the 3D spatial relationship between the human body and the device; if the projection node coordinates are on the device... Within the area of the spatial plane bound to the candidate node, the distance from the candidate human node to the spatial plane bound to the candidate device node is calculated based on the three-dimensional spatial relationship between the human body and the device, and is taken as the distance of the human key node. If the coordinates of the projected node are not within the area of the spatial plane bound to the candidate device node, the distance from the candidate human node to the spatial line segment bound to the candidate device node is calculated based on the three-dimensional spatial relationship between the human body and the device, and the point-to-line distance is obtained. The minimum value between the point-to-line distance and the candidate distance is taken as the distance of the human key node corresponding to the human key node. The long and short distances of each human key node distance are discarded according to a preset ratio to obtain the remaining human key node distances. The average value of the remaining human key node distances is calculated to obtain the distance between the person and the device.
[0108] In an optional implementation, the human body space construction module 804 is further configured to perform mask matching based on the left-eye human body segmentation result and the right-eye human body segmentation result to obtain the binocular human body key node pixel coordinates of each of the human body key nodes; for each of the human body key nodes, calculate the horizontal coordinate difference based on the left-eye human body key node pixel coordinates and the right-eye human body key node pixel coordinates in the binocular human body key node pixel coordinates to obtain the human body key node disparity; calculate the three-dimensional coordinates of the human body key nodes in the left-eye camera coordinate system based on the intrinsic parameter matrix and baseline length of the binocular camera, as well as the human body key node disparity and the left-eye human body key node pixel coordinates to obtain the three-dimensional coordinates of the human body key nodes; project the three-dimensional coordinates of each of the human body key nodes onto the left-eye camera coordinate system to obtain the spatial relationship of each of the human body key nodes in the left-eye camera coordinate system.
[0109] In an optional implementation, the device space construction module 801 is further configured to: calculate the disparity of each device key node by performing abscissa difference calculation based on the pixel coordinates of the device key node in the left eye image and the pixel coordinates of the device key node in the right eye image; calculate the three-dimensional coordinates of the device key node in the left eye camera coordinate system based on the intrinsic parameter matrix and baseline length of the binocular camera, as well as the disparity of the device key node and the pixel coordinates of the device key node in the left eye image; perform spatial plane equation calculation and spatial line segment equation calculation based on the three-dimensional coordinates of the device key node to obtain the spatial plane and spatial line segment corresponding to the device key node; and bind the three-dimensional coordinates of the device key node, the spatial plane and spatial line segment corresponding to the device key node to obtain the spatial relationship of the device key node in the left eye camera coordinate system.
[0110] In an optional implementation, the current binocular behavior detection result includes left-eye behavior detection result and right-eye behavior detection result. The hazard behavior detection module 802 is further configured to input the left-eye image of the substation in the current binocular visual image of the substation into the personnel behavior detection model to obtain the left-eye behavior detection result; if there is no hazard behavior rectangle in the left-eye behavior detection result, the step of sampling the current substation binocular visual image at intervals using the binocular camera is re-executed; if there is a hazard behavior rectangle in the left-eye behavior detection result, the right-eye image of the substation in the current binocular visual image of the substation is input into the personnel behavior detection model to obtain the right-eye behavior detection result; if there is a hazard behavior rectangle in the right-eye behavior detection result, the real-time substation binocular visual image is acquired using the binocular camera; if there is no hazard behavior rectangle in the right-eye behavior detection result, the step of sampling the current substation binocular visual image at intervals using the binocular camera is re-executed.
[0111] In an optional implementation, the preset distance indicators include a first danger distance, a second danger distance, a third danger distance, and a fourth danger distance. The hazard difference alarm module 806 is further configured to: if the distance between the personnel and the equipment is less than the first danger distance, then the hazard alarm level is a first-level alarm; if the distance between the personnel and the equipment is less than the second danger distance, then the hazard alarm level is a second-level alarm; if the distance between the personnel and the equipment is less than the third danger distance, then the hazard alarm level is a first-level warning; if the distance between the personnel and the equipment is less than the fourth danger distance, then the hazard alarm level is a second-level warning; if the distance between the personnel and the equipment is greater than or equal to the fourth danger distance, then no warning is issued.
[0112] In an optional implementation, the personnel behavior detection model is an improved YOLO12-m target detection model. The improved YOLO12-m target detection model includes a backbone network, a feature pyramid network, and a head detection network. The feature pyramid network includes a first-scale interactive attention module, a second-scale interactive attention module, a first convolutional layer, a second convolutional layer, a first efficient convolutional attention module, a second efficient convolutional attention module, and a third efficient convolutional attention module. The head detection network includes a first classification regression module, a second classification regression module, and a third classification regression module. The hidden danger behavior detection module 802 is further used to extract features from the left eye image of the substation through the backbone network to obtain a first-layer feature map, a second-layer feature map, and a third-layer feature map. The third-layer feature map and the second-layer feature map are input into the first-scale interactive attention module to obtain a first-scale interactive attention feature map, and the first-scale interactive attention feature map and the second-layer feature map are concatenated to obtain a first fusion map. The first fusion map and the first-layer feature map are input into the second-scale interactive attention module to obtain a second-scale interactive attention feature map, and the second-scale interactive attention feature map and the first-layer feature map are concatenated. The first convolutional layer is concatenated to obtain a second fused image; the second fused image is input into the first convolutional layer to obtain a first feature map, and the first feature map is concatenated with the first fused image to obtain a third fused image; the third fused image is input into the second convolutional layer to obtain a second feature map, and the second feature map is concatenated with the third layer feature map to obtain a fourth fused image; the second fused image is input into the first efficient convolutional attention module to obtain a first efficient convolutional attention feature map, the third fused image is input into the second efficient convolutional attention module to obtain a second efficient convolutional attention feature map, and the fourth fused image is input into the third efficient convolutional attention module to obtain a third efficient convolutional attention feature map; the first efficient convolutional attention feature map, the second efficient convolutional attention feature map, and the third efficient convolutional attention feature map are respectively input into the first classification regression module, the second classification regression module, and the third classification regression module to obtain a first initial monocular behavior detection result, a second initial monocular behavior detection result, and a third initial monocular behavior detection result; the left-eye behavior detection result is obtained based on the first initial monocular behavior detection result, the second initial monocular behavior detection result, and the third initial monocular behavior detection result.
[0113] The apparatus provided in this embodiment can execute the steps of the substation binocular visual hazard behavior alarm method provided in Embodiment 1. To avoid repetition, it will not be described again.
[0114] The substation binocular vision-based hazard behavior alarm system proposed in this embodiment acquires binocular images of key equipment nodes using a binocular camera, thereby constructing the spatial relationships between these key nodes. By periodically sampling the substation's binocular vision images, binocular hazard behavior bounding boxes are detected, and real-time substation binocular vision images are acquired. Using these real-time images, the system detects the bounding boxes and then constructs the spatial relationships between key human body nodes using binocular human body segmentation and key point matching. Based on the spatial relationships between these key human body nodes and key equipment nodes, a sparse point cloud ranging algorithm is used to calculate the distance between personnel and equipment, and differentiated alarms are issued for different hazard alarm levels accordingly. In this way, by establishing the spatial relationship between key human body nodes and key equipment nodes in a unified left-eye camera coordinate system, the system quantifies the distance between hazard behavior and substation equipment, achieving differentiated alarms for substation hazard behavior and effectively improving the intelligence level and response efficiency of operational safety monitoring.
[0115] Example 3 Furthermore, this disclosure provides a computer device including a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the substation binocular visual hazard behavior alarm method described in Embodiment 1.
[0116] The device provided in this embodiment can execute the steps of the substation binocular vision hazard behavior alarm method provided in Embodiment 1. To avoid repetition, it will not be described again.
[0117] Example 4 This disclosure provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the substation binocular visual hazard behavior alarm method described in Embodiment 1.
[0118] In this embodiment, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0119] The computer-readable storage medium provided in this embodiment can implement the substation binocular vision hazard behavior alarm method provided in Embodiment 1. To avoid repetition, it will not be described again here.
[0120] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not as limitations; therefore, other examples of exemplary embodiments may have different values.
[0121] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0122] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A binocular visual hazard behavior alarm method for substations, characterized in that, The method includes: The binocular images of multiple key equipment nodes in a substation are acquired using a binocular camera, and the spatial relationship of each key equipment node in the coordinate system of the left camera is constructed based on the binocular images. The binocular camera is used to sample the current substation binocular visual image at intervals. The current substation binocular visual image is input into the personnel behavior detection model to obtain the current binocular behavior detection result. If there are hidden danger behavior rectangles in the current binocular behavior detection result, the binocular camera is used to collect the real-time substation binocular visual image. The real-time binocular visual image of the substation is input into the personnel behavior detection model to obtain the real-time binocular behavior detection result. If the real-time binocular behavior detection result contains the hazard behavior rectangle, the real-time binocular visual image of the substation is input into the personnel entity segmentation model to obtain the left eye human body segmentation result and the right eye human body segmentation result. Based on the human body segmentation results of the left and right eyes, human body key nodes are matched to obtain the three-dimensional coordinates of multiple human body key nodes. Based on the three-dimensional coordinates of each human body key node, the spatial relationship of each human body key node in the coordinate system of the left eye camera is constructed. Based on the spatial relationship of each key human node in the coordinate system of the left eye camera and the spatial relationship of each key device node in the coordinate system of the left eye camera, three-dimensional sparse point cloud ranging is performed to obtain the distance between the person and the device. The hazard alarm level is determined based on the distance between personnel and equipment and the preset distance index, and the hazard behavior of the substation is given differentiated alarms based on the hazard alarm level.
2. The substation binocular visual hazard behavior alarm method according to claim 1, characterized in that, The step of performing three-dimensional sparse point cloud ranging based on the spatial relationships of each key human node in the left eye camera coordinate system and the spatial relationships of each key device node in the left eye camera coordinate system to obtain the distance between the person and the device includes: Based on the spatial relationships of each key human body node in the left eye camera coordinate system and the spatial relationships of each key device node in the left eye camera coordinate system, the three-dimensional spatial relationship between the human body and the device is obtained. For each of the aforementioned key human body nodes, based on the three-dimensional spatial relationship between the human body and the device, the distances from each key human body node to all the key device nodes are calculated to obtain multiple node distances; The minimum value among the multiple node distances is taken as the candidate distance, and the human body key node and the device key node corresponding to the candidate distance are marked as human body candidate node and device candidate node, respectively. Based on the three-dimensional spatial relationship between the human body and the device, calculate the projection node coordinates from the human body candidate node to the spatial plane bound to the device candidate node; If the coordinates of the projected node are within the area of the spatial plane to which the device candidate node is bound, then based on the three-dimensional spatial relationship between the human body and the device, the distance from the human body candidate node to the spatial plane to which the device candidate node is bound is calculated and used as the distance of the human body key node; If the coordinates of the projection node are not within the area of the spatial plane to which the device candidate node is bound, then based on the three-dimensional spatial relationship between the human body and the device, the distance from the human body candidate node to the spatial line segment to which the device candidate node is bound is calculated to obtain the point-to-line distance. The minimum value between the point-line distance and the candidate distance is taken as the distance of the human body key node corresponding to the human body key node. According to a preset ratio, the long and short distances of each of the key human body node distances are discarded to obtain the remaining key human body node distances. The average of the remaining key human body node distances is calculated to obtain the distance between the person and the equipment.
3. The substation binocular visual hazard behavior alarm method according to claim 1, characterized in that, The step of matching key human body nodes based on the left-eye and right-eye human body segmentation results to obtain the three-dimensional coordinates of multiple key human body nodes, and constructing the spatial relationship of each key human body node in the left-eye camera coordinate system based on the three-dimensional coordinates of each key human body node, includes: Based on the left-eye human body segmentation result and the right-eye human body segmentation result, mask matching is performed to obtain the binocular human body key node pixel coordinates of each of the human body key nodes; For each of the aforementioned key human body nodes, the horizontal coordinate difference is calculated based on the left and right eye key human body node pixel coordinates in the binocular key human body node pixel coordinates to obtain the disparity of the key human body node. Based on the intrinsic parameter matrix and baseline length of the binocular camera, as well as the disparity of the human body key node and the pixel coordinates of the left human body key node, the three-dimensional coordinates of the human body key node in the coordinate system of the left camera are calculated to obtain the three-dimensional coordinates of the human body key node. The three-dimensional coordinates of each of the key human body nodes are projected onto the left eye camera coordinate system to obtain the spatial relationship of each of the key human body nodes in the left eye camera coordinate system.
4. The substation binocular visual hazard behavior alarm method according to claim 1, characterized in that, The device's binocular images include a left-eye image and a right-eye image. Constructing the spatial relationships of each key node of the device in the left-eye camera coordinate system based on the binocular images includes: For each of the key nodes of the device, the parallax of the key node is obtained by calculating the difference in the horizontal coordinates between the pixel coordinates of the key node in the left eye image of the device and the pixel coordinates of the key node in the right eye image of the device. Based on the intrinsic parameter matrix and baseline length of the binocular camera, the disparity of the key nodes of the device and the pixel coordinates of the key nodes of the device in the left eye image of the device, the three-dimensional coordinates of the key nodes of the device in the coordinate system of the left eye camera are calculated to obtain the three-dimensional coordinates of the key nodes of the device. Based on the three-dimensional coordinates of the key nodes of the equipment, spatial plane equations and spatial line segment equations are calculated to obtain the spatial planes and spatial line segments corresponding to the key nodes of the equipment. The spatial relationship of the key nodes of the device in the left eye camera coordinate system is obtained by binding the three-dimensional coordinates of the key nodes of the device, the spatial planes and spatial line segments corresponding to the key nodes of the device.
5. The substation binocular visual hazard behavior alarm method according to claim 1, characterized in that, The current binocular behavior detection results include left-eye behavior detection results and right-eye behavior detection results. The current binocular visual image of the substation is input into the personnel behavior detection model to obtain the current binocular behavior detection results. If all current binocular behavior detection results contain hazard behavior rectangles, then the binocular camera is used to acquire real-time binocular visual images of the substation, including: The left-eye image of the substation in the current binocular vision image of the substation is input into the personnel behavior detection model to obtain the left-eye behavior detection result; If no hidden danger behavior rectangle is found in the left eye behavior detection result, the step of sampling the current substation binocular visual image at intervals using the binocular camera is repeated. If the left-eye behavior detection result contains the hidden danger behavior rectangle, then the right-eye image of the substation in the current binocular visual image of the substation is input into the personnel behavior detection model to obtain the right-eye behavior detection result; If the right eye behavior detection result contains the rectangle of the hidden danger behavior, then the binocular camera is used to acquire the real-time binocular visual image of the substation. If the right eye behavior detection result does not contain the rectangle of the hidden danger behavior, then the step of sampling the current substation binocular visual image at intervals using the binocular camera is repeated.
6. The substation binocular visual hazard behavior alarm method according to claim 1, characterized in that, The preset distance indicators include a first danger distance, a second danger distance, a third danger distance, and a fourth danger distance. Determining the hazard alarm level based on the personnel-equipment distance and the preset distance indicators includes: If the distance between the personnel and the equipment is less than the first danger distance, then the hazard alarm level is a first-level alarm. If the distance between the personnel and the equipment is less than the second danger distance, then the hazard alarm level is a second-level alarm. If the distance between the personnel and the equipment is less than the third danger distance, then the hazard alarm level is the first level warning. If the distance between the personnel and the equipment is less than the fourth danger distance, the hazard alarm level is a second-level warning. If the distance between the personnel and the equipment is greater than or equal to the fourth danger distance, no warning will be given.
7. The substation binocular visual hazard behavior alarm method according to claim 5, characterized in that, The personnel behavior detection model is an improved YOLO12-m object detection model. The improved YOLO12-m object detection model includes a backbone network, a feature pyramid network, and a head detection network. The feature pyramid network includes a first-scale interactive attention module, a second-scale interactive attention module, a first convolutional layer, a second convolutional layer, a first efficient convolutional attention module, a second efficient convolutional attention module, and a third efficient convolutional attention module. The head detection network includes a first classification regression module, a second classification regression module, and a third classification regression module. The step of inputting the left-eye image of the substation from the current binocular visual image of the substation into the personnel behavior detection model to obtain the left-eye behavior detection result includes: The backbone network is used to extract features from the left eye image of the substation to obtain a first layer feature map, a second layer feature map and a third layer feature map. The third-layer feature map and the second-layer feature map are input into the first-scale interactive attention module to obtain the first-scale interactive attention feature map. The first-scale interactive attention feature map and the second-layer feature map are concatenated to obtain the first fusion map. The first fused image and the first layer feature map are input into the second scale interactive attention module to obtain the second scale interactive attention feature map, and the second scale interactive attention feature map is concatenated with the first layer feature map to obtain the second fused image; The second fusion map is input into the first convolutional layer to obtain the first feature map, and the first feature map is concatenated with the first fusion map to obtain the third fusion map; The third fusion map is input into the second convolutional layer to obtain the second feature map, and the second feature map is concatenated with the third layer feature map to obtain the fourth fusion map; The second fused image is input into the first efficient convolutional attention module to obtain the first efficient convolutional attention feature map; the third fused image is input into the second efficient convolutional attention module to obtain the second efficient convolutional attention feature map; and the fourth fused image is input into the third efficient convolutional attention module to obtain the third efficient convolutional attention feature map. The first efficient convolutional attention feature map, the second efficient convolutional attention feature map, and the third efficient convolutional attention feature map are respectively input into the first classification and regression module, the second classification and regression module, and the third classification and regression module to obtain the first initial monocular behavior detection result, the second initial monocular behavior detection result, and the third initial monocular behavior detection result; The left-eye behavior detection result is obtained based on the first initial monocular behavior detection result, the second initial monocular behavior detection result, and the third initial monocular behavior detection result.
8. A binocular vision-based hazard warning system for substations, characterized in that, The system includes: The equipment space construction module is used to acquire binocular images of multiple key equipment nodes in a substation through a binocular camera, and construct the spatial relationship of each key equipment node in the coordinate system of the left eye camera based on the binocular images. The hazard behavior detection module is used to periodically sample the current substation binocular visual images using the binocular camera, input the current substation binocular visual images into the personnel behavior detection model, and obtain the current binocular behavior detection results. If there are hazard behavior rectangles in the current binocular behavior detection results, then the binocular camera is used to collect real-time substation binocular visual images; the real-time substation binocular visual images are input into the personnel behavior detection model to obtain real-time binocular behavior detection results. The personnel entity segmentation module is used to input the real-time binocular visual image of the substation into the personnel entity segmentation model if the real-time binocular behavior detection results all contain the rectangle of the hidden danger behavior, so as to obtain the left-eye human body segmentation result and the right-eye human body segmentation result. The human body space construction module is used to match key human body nodes based on the left-eye human body segmentation results and the right-eye human body segmentation results, obtain the three-dimensional coordinates of multiple key human body nodes, and construct the spatial relationship of each key human body node in the left-eye camera coordinate system based on the three-dimensional coordinates of each key human body node. The three-dimensional point cloud ranging module is used to perform three-dimensional sparse point cloud ranging based on the spatial relationship of each of the key human body nodes in the left eye camera coordinate system and the spatial relationship of each of the key equipment nodes in the left eye camera coordinate system, so as to obtain the distance between the personnel and the equipment. The hazard differentiation alarm module is used to determine the hazard alarm level based on the distance between personnel and equipment and a preset distance indicator, and to issue differentiated alarms for hazard behaviors in the substation based on the hazard alarm level.
9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the substation binocular visual hazard behavior alarm method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the substation binocular visual hazard behavior alarm method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method, system and terminal for detecting illegal behaviors of person standing under suspension arm
CN114494427A
Intelligent monitoring method for cross operation in refining device area
CN117475343A
Transformer substation safety construction monitoring and early warning system based on binocular stereoscopic vision
CN121354324A
Petroleum drilling machine equipment intelligent alarm method based on video real-time display
CN121691594A
Video surveillance system, surveillance video composition apparatus, and video surveillance server
US20040257444A1