A camera tracking method based on human skeleton points in a power distribution substation

By using a camera tracking method based on human skeleton points, the optimal shooting position and angle of the moving camera are calculated, which solves the problems of easy obstruction of monitoring images, insufficient resolution, and insufficient early warning in the substation monitoring system, and realizes efficient tracking of maintenance personnel and early warning of dangerous actions.

CN118505744BActive Publication Date: 2026-05-26BEIJING INFORMATION SCI & TECH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INFORMATION SCI & TECH UNIV
Filing Date
2024-05-11
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing substation monitoring systems suffer from several problems, including easily obstructed monitoring images leading to high false alarm rates, insufficient resolution for long-distance shooting, inability to provide early warnings of dangerous actions, and the inability to dynamically adjust the warning range of electronic fences.

Method used

A camera tracking method based on human skeleton points is adopted. By calculating the optimal shooting angle and position of the moving camera, the position adjustment of the moving camera is guided by the fixed camera. The extrinsic parameter matrix is ​​optimized by combining particle swarm optimization algorithm, and the position and operation of maintenance personnel are determined by deep learning and skeleton point recognition algorithm. Kalman filtering is used for early warning.

Benefits of technology

It solves the problem of false alarms caused by obstructed monitoring images, improves the resolution and recognition rate of long-distance shooting, realizes real-time tracking of maintenance personnel and early warning of dangerous actions, and dynamically adjusts the warning range of electronic fences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118505744B_ABST
    Figure CN118505744B_ABST
Patent Text Reader

Abstract

This disclosure provides a camera tracking method based on human skeleton points in a power distribution substation, belonging to the field of power distribution substation operation and maintenance. The method includes constructing an operation and maintenance world coordinate system, fixing the camera to identify the personnel's position, calculating the optimal shooting angle and position of the moving camera, moving the moving camera to the optimal position, determining whether the personnel have crossed an electronic fence, and determining whether an incorrect operation component has been pressed. By using a fixed camera to guide the moving camera to the optimal monitoring angle and position where the operation and maintenance personnel can be filmed, it solves the problem of high false alarm rates caused by obstructed individual monitoring images, the problem of decreased recognition rate due to low resolution from long-distance shooting, the problem of determining whether the operation and maintenance personnel have crossed an electronic fence based on the security level, and the problem of determining whether the operation and maintenance personnel are about to operate an operation component in a specific operation and maintenance task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power distribution station operation and maintenance technology. Background Technology

[0002] Traditional power distribution station operation and maintenance (O&M) monitoring technologies primarily rely on on-site manual inspections and remote manual monitoring via cameras, lacking automated and intelligent safety management methods. Today, intelligent O&M monitoring technologies for power distribution stations benefit from advancements in network technology, deep learning, and machine vision. The expanding coverage of surveillance videos and the continuously improving clarity of monitoring images, coupled with high-speed network transmission ensuring accurate and rapid information transmission, enable deep learning and machine vision technologies to perform more efficient and accurate analysis of the images.

[0003] The video surveillance scope of power grid sites mainly includes the following three aspects: equipment status, personnel status, and environmental status of important areas of the power grid. For the first aspect, equipment status, there are two parts. Firstly, transmission line equipment, such as whether insulators, transmission towers, and transmission lines are aging, and whether there is snow accumulation on the transmission lines. Secondly, equipment within the power station, such as whether the equipment exterior is damaged, the status of disconnect switches, circuit breakers, instrument readings, and whether important warning signs are displayed in designated locations. For the second aspect, personnel status, the focus is on identifying personnel and whether their attire meets the safety level of the current working environment. For the third aspect, the environmental status of important areas of the power grid, the focus is on whether there are potential natural and man-made risks in the areas where equipment is located, such as smoke sources, open flames, the number of personnel in insulated areas, and foreign object intrusion.

[0004] However, existing intelligent monitoring systems also have many problems: 1. Existing monitoring systems mainly use on-site PTZ cameras and fixed network cameras. Individual monitoring images are easily obstructed, leading to a high rate of false alarms. 2. Fixed camera installation locations cannot be moved, and large substations have large working areas, making long-distance shooting resolution insufficient. 3. They can only record after dangerous actions occur, and cannot truly provide early warnings of dangerous actions. 4. They cannot dynamically adjust the warning range of the electronic fence according to the high-voltage control level. Summary of the Invention

[0005] To overcome the above-mentioned defects of the prior art, the present invention provides a camera tracking method for substation operation and maintenance personnel based on human skeleton points.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] This invention provides a method for calculating the optimal shooting angle and position for a moving camera. This method utilizes images captured by a fixed camera, with the following constraints: the three-dimensional coordinates of the center point of the target frame, the three-dimensional coordinates of the center points of other personnel target frames, the radius of the simulated body, the redundant viewing distance of the binocular camera, the limit angle of the moving binocular camera, the three-dimensional coordinates of the spatial position of the operating component, and the fixed installation height of the lateral slide rail relative to the origin. The output parameter is the extrinsic parameter matrix of the moving binocular camera. Optimization is performed using a particle swarm optimization algorithm.

[0008] Another aspect of the present invention provides a method for camera tracking of human skeleton points by substation maintenance personnel, the method comprising the following steps:

[0009] (1) Construct an operation and maintenance world coordinate system;

[0010] (2) Fixed camera identifies the location of people;

[0011] (3) Calculate the best shooting position for the moving camera;

[0012] (4) Move the camera to the optimal position;

[0013] (5) Determine whether the electronic fence has been crossed;

[0014] (6) Determine whether to operate the operating component.

[0015] The specific steps included in the camera tracking method for substation maintenance personnel based on human skeleton points are as follows:

[0016] (1) Constructing the Operation and Maintenance World Coordinate System

[0017] The intrinsic and extrinsic parameter matrices of the fixed camera are determined, the intrinsic parameter matrix of the moving camera is determined, the origin of the world coordinate system is determined, reference markers are placed on the boundary of the electronic fence, reference markers are placed on the position of the operating component, the positions of the electronic fence and the operating component in the world coordinate system are calibrated, and the geometry of the operating component is generated.

[0018] (2) Fixed camera identifies personnel position

[0019] The scene is captured using a calibrated binocular camera, and a deep learning algorithm is used to identify maintenance personnel and calculate their absolute three-dimensional coordinates.

[0020] (3) Calculate the best shooting position for the moving camera

[0021] The extrinsic parameter matrix of the moving binocular camera was calculated using image processing and spatial geometric constraint methods.

[0022] (4) Move the camera to the best position.

[0023] Using a two-degree-of-freedom "H"-shaped sliding rail, the camera is moved to a designated position based on the extrinsic and intrinsic parameter matrices, and the moving binocular camera is used to capture real-time images of the maintenance personnel.

[0024] (5) Determine whether the electronic fence has been crossed

[0025] The skeleton point recognition algorithm calculates the three-dimensional coordinates of the skeleton points of maintenance personnel and determines whether the skeleton points are within the range of the electronic fence based on the prevention and control level.

[0026] (6) Determine whether to operate the operating component.

[0027] Save the skeletal coordinates of the maintenance personnel in the time matrix, use Kalman filtering to predict the skeletal coordinates of the maintenance personnel at the next moment, and determine whether the predicted coordinates are within the geometry of the operating component.

[0028] Based on the above, this disclosure describes a camera tracking method for substation maintenance personnel based on human skeleton points. By using a fixed camera to guide the moving camera to the optimal monitoring angle and position where the maintenance personnel can be filmed, it solves the problems of high false alarm rates caused by obstructed individual monitoring images, low recognition rates due to low resolution from long-distance shooting, determining whether maintenance personnel have crossed electronic fences based on security levels, and determining whether maintenance personnel are about to operate components during special maintenance tasks. Attached Figure Description

[0029] Figure 1 Spatial schematic diagram of a multi-person motion-sensing image detection system in a power distribution station;

[0030] Figure 2 Top view schematic diagram of a multi-person motion-sensing image detection system for a power distribution station;

[0031] Figure 3 A diagram illustrating the movement of a homing camera to avoid obstructing human bodies and to find the best viewing angle for visible operating parts;

[0032] Figure 4 Schematic diagram of the operation process of a multi-person motion-sensor image detection system in a power distribution station;

[0033] In the diagram: 1. Fixed camera; 2. Moving camera; 3. Horizontal slide rail; 4. Vertical follow-up slide rail; 5. Power distribution cabinet; 6. Maintenance personnel; 7. Insulated area; 8. Operating components. Detailed Implementation

[0034] The specific implementation methods of this invention are further described below. Except for the content specifically mentioned below, the processes, conditions, and experimental methods for implementing this invention are all common knowledge and general knowledge in the field. This invention does not have any particular limitations. Spatial diagrams and top views are as follows. Figure 1 and Figure 2 As shown.

[0035] This invention provides a camera tracking method for substation maintenance personnel based on human skeleton points, which detects whether maintenance personnel have crossed electronic fences and whether they are about to press the wrong operating device. Its workflow diagram is shown below. Figure 4 As shown:

[0036] (1) Constructing the Operation and Maintenance World Coordinate System

[0037] according to Figure 1 and Figure 2 The following describes how to set up a monitoring world environment:

[0038] (a) Use Zhang Zhengyou's calibration method to calibrate the intrinsic and extrinsic parameter matrices of fixed cameras A and B;

[0039] (b) Select a special location in a room as the origin of the world coordinate system. This special location should be easy to identify and locate, such as a corner of the room. Place a fixed reference mark at this location as the selected origin. This mark will be used for subsequent camera calibration and coordinate transformation. Use the calibrated camera to take an image with the reference mark. Use the intrinsic and extrinsic parameter matrices of fixed cameras A and B to determine that point as the origin of the world coordinate system.

[0040] (c) Place a fixed reference mark at the position of the operating part, and use calibrated fixed cameras A and B to take images with the reference mark. The reference mark will be used for subsequent camera calibration and coordinate transformation. Generate a geometry of the operating part in the computer terminal based on the transformed three-dimensional coordinates. For example, if the operating part is a button, generate a sphere μ with a radius of 3cm with the three-dimensional coordinate E of the button as the origin. The position of the button in the world coordinate system is determined by this sphere.

[0041] (d) Place multiple fixed reference markers on the boundary of the electronic fence, and use calibrated fixed cameras A and B to take images with the reference markers. The reference markers will be used for subsequent camera calibration and coordinate transformation. The range of the electronic fence is defined in the computer terminal based on the transformed three-dimensional coordinates.

[0042] (e) Place two moving cameras on each transverse follower slide rail, place a calibration plate at a certain coordinate position, keep the shooting angle of the camera unchanged, move the camera once at a certain distance, obtain the camera extrinsic parameter matrix for each movement, and calculate the straight line equation of the slide rail in the world coordinate system.

[0043] (f) Place two moving cameras on each longitudinal follower slide rail, place a calibration plate at a certain coordinate position, keep the shooting angle of the camera unchanged, move the camera once at a certain distance, obtain the camera extrinsic parameter matrix for each movement, and calculate the straight line equation of the slide rail in the world coordinate system.

[0044] (2) Fixed camera identifies personnel position

[0045] Fixed cameras identify the 3D coordinates of maintenance personnel in the world coordinate system. Calibrated fixed multi-view cameras A and B capture images of the tested scene and transmit them via fiber optic cable to a computer terminal. Deep learning algorithms are used to identify rectangular frames representing human figures, with the center point of each frame considered as the absolute 3D coordinates P of n maintenance personnel in the world coordinate system. N (N = 1:n).

[0046] (3) Calculate the best shooting position for the moving camera

[0047] For a moving camera, the algorithm calculation for camera-guided avoidance of occlusion and human movement is illustrated in the diagram below. Figure 3 As shown; the constraints include:

[0048] (a) The three-dimensional coordinates p1 of the center point of the target frame of the target maintenance personnel in the world coordinate system;

[0049] (b) The three-dimensional coordinates of the center point of the target frame of the other n-1 maintenance personnel in the world coordinate system: P2, P3...P n ;

[0050] (c) The radius r of the cylinder containing the body of the maintenance personnel;

[0051] (d) The generated three-dimensional coordinates E of the operating component space;

[0052] (e) The limiting angle α of the field of view of the moving binocular camera;

[0053] (f) The optimal parallax angle β for binocular camera observation;

[0054] (g) The equation of the spatial straight line γ relative to the origin of the world coordinate system obtained after the horizontal slide rail is calibrated;

[0055] (h) The spatial straight line equations δ1, δ2, δ3, and δ4 relative to the origin of the world coordinate system obtained after the longitudinal slide rail calibration are given by the constraint relationship formulas of the extrinsic parameter matrices of the m moving cameras:

[0056] (R m , t m )=f(P N , r, E, α, β, γ1, γ2, δ1, δ2, δ3, δ4), (N=1:n)

[0057] The particle swarm optimization algorithm is used to find the optimal solution of the function, where:

[0058] (a) Initial particle count num_particles = 30;

[0059] (b) Maximum number of iterations max_iterations = 100;

[0060] (c) Velocity range = 0.1;

[0061] (d) Acceleration constant factors = [0.5, 1.5];

[0062] (e) Inertia weight = 0.9;

[0063] The solution process is performed on the computing terminal.

[0064] (4) Move the camera to the best position.

[0065] Based on the extrinsic parameter matrix and known intrinsic parameter matrix of the binocular moving camera obtained in (3), the "H" shaped slide rail is used to move the moving camera to the optimal shooting position. The two-degree-of-freedom "H" shaped slide rail consists of a horizontal follower slide rail and a vertical follower slide rail. The horizontal follower slide rail can slide on the vertical follower slide rail, and the moving binocular camera slides on the horizontal follower slide rail. A distance sensor is installed at the connection point, and the moving distance of the slide rail is controlled by the command issued by the computing terminal.

[0066] (5) Determine whether the electronic fence has been crossed

[0067] After the mobile binocular camera reaches the designated location and captures images, skeletal point recognition algorithms, such as OpenPose and YOLO-Pose, are used to calculate the spatial three-dimensional coordinates {x(i)} of the maintenance personnel's skeletal points based on the mobile camera's intrinsic and extrinsic parameter matrices. n ,y(i) n , z(i) n}, where i = 1:17 represents the key point number, n = 1:N represents the number of the personnel being tested in the operation and maintenance scenario; model in the computer terminal, determine whether the three-dimensional coordinates of each bone point are within the range of the electronic fence established in (1), and combine different prevention and control levels to determine whether the three-dimensional coordinates of each bone point are close to the edge of the electronic fence range established in (1), and issue an alarm.

[0068] (6) Determine whether to operate the erroneous component.

[0069] The spatial three-dimensional coordinates of the operation and maintenance personnel's skeletal points are stored in a time matrix. This time matrix stores the spatial three-dimensional coordinates at the current moment. It also preserves the spatial three-dimensional coordinates from the previous moment. The "moment" mentioned here refers to the frame number of the video, specifically the spatial 3D coordinates of the skeleton point of the maintenance personnel in the previous frame. The Kalman filter method is used to predict the 3D spatial coordinates of the skeleton point in the next moment based on the coordinate matrix in the time dimension. like If an early warning is issued within the operating component body μ established in (1), then an early warning will be issued.

Claims

1. A camera tracking method based on human skeleton points in a power distribution station, characterized in that, Includes the following steps: Both fixed and mobile multi-view cameras were used. The fixed multi-view camera was used to identify the absolute three-dimensional coordinates of the maintenance personnel in the world coordinate system and to calculate the best shooting position that the mobile multi-view camera should take. The camera extrinsic parameter matrix of the mobile multi-view camera is calculated; after the mobile multi-view camera arrives at the location according to the instructions, it begins to collect images of the maintenance personnel; it determines whether the maintenance personnel have touched the electronic fence and issues a real-time warning; It can detect whether maintenance personnel are about to touch the wrong operating component and provide real-time warnings. Construct a world coordinate system for substation operation and maintenance: determine the intrinsic and extrinsic parameter matrices of fixed multi-cameras, determine the intrinsic parameter matrix of mobile multi-cameras, and determine the origin of the world coordinate system; place reference markers on the boundaries of the electronic fence, place reference markers on the positions of operating components, calibrate the positions of the electronic fence and operating components in the world coordinate system, and generate the geometry of the operating components; Calculate the optimal shooting position for the mobile camera: using a constraint method based on image processing and spatial geometry, the constraints include: the three-dimensional coordinates of the target maintenance personnel. Other maintenance personnel's three-dimensional coordinate set The following parameters are considered: the radius r of the cylindrical model representing the space occupied by the human body; the three-dimensional coordinates E of the generated operating component; the viewing angle limit α and the optimal parallax angle β of the moving multi-view camera; and the spatial straight line equation relative to the origin of the world coordinate system obtained after the lateral sliding rail calibration. The spatial straight line equation obtained after longitudinal slide rail calibration relative to the origin of the world coordinate system Therefore, the constraint relationship formula for the extrinsic parameter matrix of each moving camera is obtained as follows: The particle swarm optimization algorithm is used to find the optimal solution of the function; the obtained extrinsic matrix and the known camera intrinsic matrix will be used in subsequent image processing and 3D reconstruction tasks. The mobile binocular camera reaches a position according to instructions. Its feature is that it uses a two-degree-of-freedom "H"-shaped slide rail, which consists of a horizontal follower slide rail and a vertical follower slide rail. The horizontal follower slide rail can slide on the vertical follower slide rail, and the mobile binocular camera slides on the horizontal follower slide rail. The method for determining whether maintenance personnel have touched the electronic fence is characterized in that after the mobile binocular camera reaches the designated location and takes pictures, a skeletal point recognition algorithm is used to calculate the spatial three-dimensional coordinates of the maintenance personnel's skeletal points based on the camera's intrinsic and extrinsic parameter matrix, and the three-dimensional coordinates of each skeletal point are determined based on spatial geometry principles to determine whether the three-dimensional coordinates of each skeletal point are within the range of the electronic fence. The aforementioned early warning method for determining whether maintenance personnel are about to touch an erroneous operation component is characterized by storing the spatial three-dimensional coordinates of the maintenance personnel's skeletal points in a time matrix. This time matrix stores not only the spatial three-dimensional coordinates at the current moment but also the spatial three-dimensional coordinates at the previous moment. Based on the coordinate matrix in the time dimension, the Kalman filter method is used to predict the three-dimensional spatial coordinates of the skeletal points at the next moment. Predict the three-dimensional spatial coordinates of the skeletal point at the next moment based on the time matrix, and determine whether the coordinates are located within the geometry of the operating component.

2. The camera tracking method based on human skeleton points in a substation according to claim 1, characterized in that, The method of using a fixed camera to identify the absolute three-dimensional coordinates of maintenance personnel in the world coordinate system involves using a calibrated multi-view camera to collect real-time data on the scene under test, and using a deep learning algorithm to identify the rectangular frame of the human body. The center point of the rectangular frame is regarded as the absolute three-dimensional coordinates of the maintenance personnel in the world coordinate system.

3. The camera tracking method based on human skeleton points in a substation according to claim 1, characterized in that, The particle swarm optimization algorithm is used to find the optimal solution of the function, where: the initial number of particles num_particles = 30, the maximum number of iterations max_iterations = 100, the velocity range velocity_range = 0.1, the acceleration constant factor acceleration_factors = [0.5, 1.5], and the inertia weight inertia_weight = 0.

9.

4. The method for determining an electronic fence based on human skeleton points in a substation according to claim 1, characterized in that... Multiple fixed reference markers are placed on the boundary of the electronic fence, and images with the reference markers are taken using a calibrated camera. The reference markers will be used for subsequent camera calibration and coordinate transformation. The range of the electronic fence is defined in the computer terminal based on the transformed three-dimensional coordinates of the reference markers.

5. The camera tracking method based on human skeleton points in a substation according to claim 1, and the method for determining the position of operating components, are characterized in that... A certain number of fixed reference markers are placed at the position of the operating component, and images with the reference markers are captured using a calibrated camera. The reference markers will be used for subsequent camera calibration and coordinate transformation. Based on the transformed 3D coordinates, a geometry of the operating component is generated in the computer terminal. The operating component is a button, and a sphere with a radius of 3cm is generated with the button's 3D coordinates E as the origin. The sphere is used to determine the spatial position of the button in the world coordinate system.