A power grid aerial operation target tracking method and device, a control ball and a storage medium
Patent Information
- Application Number
- CN202610905031.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]本发明提供一种电网高空作业目标追踪方法、装置、布控球及存储介质,用以解决现有方法因依赖固定全局特征模板,在复杂背景、小目标、目标特征变化等场景下出现的追踪偏移、追踪丢失问题,提升电网高空作业目标追踪的精准度和稳定性,满足高精度电网高空作业安全监控的需求
[0010] The power grid high-altitude operation target tracking method provided in this invention analyzes the temporal differences of consecutive frames in a power grid high-altitude operation video stream to select candidate target image blocks containing potential targets. It then reconstructs these blocks using topological normalization of foreground pixel distribution density to obtain standard-scale image blocks with a unified geometric reference. This effectively solves the problem of difficult feature extraction caused by the low pixel ratio and inconsistent geometric scale of small targets. Based on these standard-scale image blocks, high-frequency detail enhancement and edge contour sharpening reconstruction are performed on texture gradient features to obtain enhanced target image blocks, improving the recognizability of small target features. Furthermore, by extracting skeleton node coordinates to construct a local deformation descriptor, accurate representation of the target's posture and geometric features is achieved. Finally, based on the local deformation descriptor, the rationality of the human body structure is verified through the relative topological relationships of nodes, and the target is further refined. Valid tracking targets are selected, and logical constraints are deduced by combining the spatial relationship between the valid tracking targets and the surrounding environment and the temporal changes of local deformation descriptors to obtain expected behavior constraints, avoiding feature misjudgment caused by rapid target movement or brief occlusion. A dynamic search space is constructed based on the instantaneous movement trend of the valid tracking targets and expected behavior constraints. Target tracking results are obtained through reverse geometric mapping and consistency verification, effectively avoiding problems such as trajectory breakage, identity switching, false detection, and missed detection, and achieving accurate target location locking. Ultimately, this solves the tracking offset and tracking loss problems that occur in existing methods due to reliance on fixed global feature templates in scenarios such as complex backgrounds, small targets, and changing target features. This improves the accuracy and stability of target tracking in high-altitude power grid operations and meets the needs of high-precision safety monitoring of high-altitude power grid operations.
Smart Images

Figure CN122780337A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method, device, control ball, and storage medium for tracking targets during high-altitude operations in power grids. Background Technology
[0002] With the continuous expansion of power grid construction, the workload of high-altitude maintenance of power grid facilities such as transmission line towers and substations is constantly increasing, and high-altitude operations in the power grid are becoming more frequent. Real-time target tracking of personnel, tools, and equipment is a crucial aspect of ensuring operational safety and efficiency. Existing target tracking methods for high-altitude power grid operations mainly rely on global feature matching of image frames. This involves extracting global image features from the current frame and performing similarity matching with a pre-defined fixed global feature template. Based on the matching result, the target's position in the current frame is determined, thus enabling continuous target tracking.
[0003] However, in actual power grid high-altitude operations, the complex and ever-changing background environment often causes the target to exhibit small characteristics. For example, dense transmission lines and complex tower structures can easily interfere with the target. Dramatic fluctuations in light intensity and extremely low pixel ratios due to long-distance shooting can all affect the recognition of target features. At the same time, the target may move rapidly, be briefly occluded, or overlap with similar background textures during the operation, causing significant changes in the target's global appearance features. This makes it difficult for existing matching methods based on fixed global feature templates to reliably identify target features, leading to problems such as trajectory breaks, identity switching, or false positives and false negatives. It is impossible to accurately locate the target position, resulting in tracking deviations and tracking loss, ultimately failing to meet the high-precision requirements for safety monitoring of power grid high-altitude operations. Summary of the Invention
[0004] This invention provides a method, device, control ball, and storage medium for tracking targets in high-altitude power grid operations. It addresses the problems of tracking offset and tracking loss that occur in existing methods due to their reliance on fixed global feature templates in scenarios with complex backgrounds, small targets, and changing target features. This invention improves the accuracy and stability of target tracking in high-altitude power grid operations, meeting the needs of high-precision safety monitoring of high-altitude power grid operations.
[0005] In a first aspect, the present invention provides a method for tracking targets during high-altitude operations in power grids, comprising: Temporal difference analysis was performed on continuous frame images of high-altitude power grid operation video stream to obtain candidate target image patches containing potential targets. Based on the foreground pixel distribution density of the candidate target image patches, topological normalization reconstruction was performed to obtain standard-scale image patches under a unified geometric benchmark. High-frequency detail enhancement and edge contour sharpening reconstruction are performed based on the texture gradient features of standard-scale image patches to obtain target enhanced image patches; Based on the skeleton node coordinates extracted from the target augmented image patch, a local deformation descriptor representing the geometric features of the target pose is constructed; The rationality of human body structure is verified based on the relative topological relationship of nodes in the local deformation descriptor, and an effective tracking target is obtained. Logical constraint deduction is then performed based on the spatial relationship between the effective tracking target and the surrounding environment, combined with the temporal changes of the local deformation descriptor, to obtain the expected behavior constraint. A dynamic search space is constructed based on the instantaneous motion trend and expected behavior constraints of the target for effective tracking. Inverse geometric mapping and consistency verification are performed in the dynamic search space to obtain the target tracking result.
[0006] In a second aspect, the present invention also provides a power grid high-altitude operation target tracking device, applied to the power grid high-altitude operation target tracking method as described in the first aspect; the power grid high-altitude operation target tracking device includes: The standardization reconstruction module is used to perform temporal difference analysis on continuous frame images of high-altitude power grid operation video stream to obtain candidate target image blocks containing potential targets, and to perform topological standardization reconstruction based on the foreground pixel distribution density of the candidate target image blocks to obtain standard-scale image blocks under a unified geometric benchmark. The target image enhancement module performs high-frequency detail enhancement and edge contour sharpening reconstruction based on the texture gradient features of standard-scale image patches to obtain target enhanced image patches; The deformation descriptor construction module is used to construct a local deformation descriptor that represents the geometric features of the target pose based on the skeleton node coordinates extracted from the target augmented image patch. The behavior constraint deduction module is used to verify the rationality of human body structure based on the relative topological relationship of nodes in the local deformation descriptor, obtain the effective tracking target, and perform logical constraint deduction based on the spatial relationship between the effective tracking target and the surrounding environment and the temporal change of the local deformation descriptor to obtain the expected behavior constraint. The dynamic search and verification module is used to construct a dynamic search space based on the instantaneous motion trend and expected behavior constraints of the effectively tracked target, and to perform reverse geometric mapping and consistency verification in the dynamic search space to obtain the target tracking result.
[0007] Thirdly, the present invention also provides a control ball, comprising: a memory for storing computer software programs; and a processor for reading and executing the computer software programs, thereby realizing the power grid high-altitude operation target tracking method as described above.
[0008] Fourthly, the present invention also provides a non-transitory computer-readable storage medium storing a computer software program, which, when executed by a processor, implements any of the above-described power grid high-altitude operation target tracking methods.
[0009] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the power grid high-altitude operation target tracking method as described above.
[0010] The power grid high-altitude operation target tracking method provided in this invention analyzes the temporal differences of consecutive frames in a power grid high-altitude operation video stream to select candidate target image blocks containing potential targets. It then reconstructs these blocks using topological normalization of foreground pixel distribution density to obtain standard-scale image blocks with a unified geometric reference. This effectively solves the problem of difficult feature extraction caused by the low pixel ratio and inconsistent geometric scale of small targets. Based on these standard-scale image blocks, high-frequency detail enhancement and edge contour sharpening reconstruction are performed on texture gradient features to obtain enhanced target image blocks, improving the recognizability of small target features. Furthermore, by extracting skeleton node coordinates to construct a local deformation descriptor, accurate representation of the target's posture and geometric features is achieved. Finally, based on the local deformation descriptor, the rationality of the human body structure is verified through the relative topological relationships of nodes, and the target is further refined. Valid tracking targets are selected, and logical constraints are deduced by combining the spatial relationship between the valid tracking targets and the surrounding environment and the temporal changes of local deformation descriptors to obtain expected behavior constraints, avoiding feature misjudgment caused by rapid target movement or brief occlusion. A dynamic search space is constructed based on the instantaneous movement trend of the valid tracking targets and expected behavior constraints. Target tracking results are obtained through reverse geometric mapping and consistency verification, effectively avoiding problems such as trajectory breakage, identity switching, false detection, and missed detection, and achieving accurate target location locking. Ultimately, this solves the tracking offset and tracking loss problems that occur in existing methods due to reliance on fixed global feature templates in scenarios such as complex backgrounds, small targets, and changing target features. This improves the accuracy and stability of target tracking in high-altitude power grid operations and meets the needs of high-precision safety monitoring of high-altitude power grid operations. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating the target tracking method for high-altitude operations in power grids provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the power grid high-altitude operation target tracking device provided in an embodiment of the present invention; Figure 3 An embodiment diagram of the control ball provided in this invention; Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0014] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0015] See Figure 1 , Figure 1 This is a flowchart illustrating the target tracking method for high-altitude operations in power grids provided by the present invention. In this embodiment, the executing entity of the target tracking method for high-altitude operations in power grids is a control sphere. Therefore, the target tracking method for high-altitude operations in power grids includes: Step 10: Perform temporal difference analysis on continuous frame images of the power grid high-altitude operation video stream to obtain candidate target image blocks containing potential targets, and perform topological normalization reconstruction based on the foreground pixel distribution density of the candidate target image blocks to obtain standard-scale image blocks under a unified geometric benchmark.
[0016] Optionally, the surveillance sphere acquires a video stream of continuous frame images in real time from built-in cameras (including but not limited to high-definition cameras and infrared cameras) in the high-altitude power grid operation scenario. It performs temporal difference analysis on the continuous frame images, wherein the acquisition time interval between two adjacent frames does not exceed 0.1 seconds. The temporal difference analysis refers to calculating the pixel value difference of each pixel in two adjacent frames, marking pixels whose pixel value difference exceeds a preset pixel difference threshold as difference pixels. All difference pixels together form a difference pixel region, which is a candidate target image block containing potential targets. The potential targets refer to the workers, tools, and equipment to be tracked in the high-altitude power grid operation scenario. The preset threshold is pre-set according to the lighting conditions of the high-altitude power grid operation scenario, with a value range of 20 to 50, and can be dynamically adjusted according to the actual scenario.
[0017] Optionally, after the control ball extracts the candidate target image patch, it calculates the distribution density of foreground pixels within the candidate target image patch. Here, foreground pixels refer to pixels within the candidate target image patch that belong to the potential target, and the foreground pixel distribution density is the ratio of the number of foreground pixels within the candidate target image patch to the total number of pixels in the candidate target image patch. Then, based on the calculated foreground pixel distribution density, the candidate target image patch is topologically normalized and reconstructed to finally obtain a standard-scale image patch under a unified geometric datum, as described in steps 101 to 105. This unified geometric datum refers to a preset fixed pixel size, pixel ratio, and pixel coordinate reference system, and the standard-scale image patch refers to an image patch that conforms to this unified geometric datum.
[0018] In one embodiment, a surveillance sphere acquires a video stream of high-altitude maintenance operations at a substation. The built-in camera has a resolution of 1920×1080 pixels, and the time interval between adjacent frame acquisitions is 0.05 seconds. The first and second consecutive frames are extracted, and the pixel value difference between the two frames is calculated pixel by pixel. A pixel difference threshold of 30 is set, and pixels with a pixel value difference greater than 30 are marked as difference pixels. The area formed by these difference pixels is a candidate target image block containing high-altitude workers.
[0019] Step 20: Based on the texture gradient features of the standard-scale image patch, perform high-frequency detail enhancement and edge contour sharpening reconstruction to obtain the target enhanced image patch.
[0020] Optionally, the control ball extracts texture gradient features from standard-scale image patches. These texture gradient features refer to the rate of change of grayscale values of pixels within the standard-scale image patch in different directions, including horizontal, vertical, 45-degree, and 135-degree angles. The rate of change of grayscale values is the ratio of the difference in grayscale values between adjacent pixels to the pixel spacing. After feature extraction, the control ball performs high-frequency detail enhancement processing on these features. High-frequency details refer to areas within the standard-scale image patch where grayscale values change rapidly, reflecting the edges and texture details of potential targets. High-frequency detail enhancement processing improves the grayscale contrast of these areas by adjusting grayscale gain, making the high-frequency detail features of the potential target more prominent. In addition to high-frequency detail enhancement, the control ball also performs edge contour sharpening reconstruction on standard-scale image blocks. The edge contour refers to the pixel grayscale value boundary area between the potential target and the background environment. Edge contour sharpening reconstruction is achieved by gradient enhancement of pixel grayscale values to make the grayscale value change of this boundary area more obvious, eliminating blurry pixels at the contour edge. Finally, the image block after high-frequency detail enhancement and edge contour sharpening reconstruction is the target enhancement image block.
[0021] Continuing with the above embodiment, assuming the standard-scale image patch is a unified geometric reference of 500×500 pixels, the control ball extracts the texture gradient features of the 500×500 pixel standard-scale image patch in step 10, calculating the grayscale value change rates horizontally, vertically, at a 45-degree angle, and at a 135-degree angle, respectively, identifying high-frequency detail areas reflecting the safety helmets and safety belts of high-altitude workers. An adaptive histogram equalization algorithm is used to adjust the grayscale value gain of this area, increasing the contrast to 1.8 times the original contrast. Subsequently, the control ball locates the boundary region between the worker's body and the background tower, enhances the pixel grayscale value gradient of this area, eliminates blurred pixels at the contour edges, completes edge contour sharpening and reconstruction, and finally obtains a target enhanced image patch containing clear worker features.
[0022] Step 30: Based on the skeleton node coordinates extracted from the target augmented image patch, construct a local deformation descriptor representing the geometric features of the target pose.
[0023] Optionally, the control ball extracts a skeleton based on the target enhancement image patch. This skeleton refers to the line connecting pixels that can characterize the overall geometric shape and structural features of the potential target, and is an abstract representation of the core geometric contour of the potential target. After the skeleton extraction is completed, key feature points on the skeleton are identified and defined as skeleton nodes. That is, skeleton nodes are the turning points, intersection points and endpoints of the geometric shape of the potential target. Different types of potential targets correspond to preset skeleton node extraction rules. For workers, skeleton nodes include key body feature points such as the head vertex, shoulder endpoints, elbow endpoints, wrist endpoints, hip midpoint, knee endpoints, and ankle endpoints.
[0024] Optionally, the control ball performs coordinate calibration on all extracted skeleton nodes to obtain the two-dimensional pixel coordinates of each skeleton node under a unified geometric reference. These two-dimensional pixel coordinates refer to the pixel position coordinates with the upper left corner of a standard-scale image patch as the origin, the horizontal axis as the x-axis, and the vertical axis as the y-axis. Based on the two-dimensional pixel coordinates of all skeleton nodes, a local deformation descriptor capable of representing the geometric features of the target's posture is constructed. This local deformation descriptor is a set of digital representations of the relative positions, distances, angles, and other geometric features between skeleton nodes, reflecting the limb posture, deformation state, and other geometric features of the potential target.
[0025] Continuing with the above embodiment, the control ball extracts the skeleton from the target augmented image block containing the high-altitude worker, obtaining skeleton lines representing the worker's body contour. According to the worker skeleton node extraction rules, 12 skeleton nodes are identified and marked: head vertex, left shoulder endpoint, right shoulder endpoint, left elbow endpoint, right elbow endpoint, left wrist endpoint, right wrist endpoint, hip midpoint, left knee endpoint, right knee endpoint, left ankle endpoint, and right ankle endpoint. These 12 skeleton nodes are then calibrated with two-dimensional pixel coordinates, obtaining specific coordinate values such as head vertex (250, 80), left shoulder endpoint (220, 150), and right shoulder endpoint (280, 150). Based on these coordinate values, geometric features such as relative distances and angles between nodes are calculated and digitally integrated to construct a local deformation descriptor reflecting the worker's limb posture while climbing the tower.
[0026] Step 40: Based on the relative topological relationship of nodes in the local deformation descriptor, the rationality of the human body structure is verified to obtain the effective tracking target. Based on the spatial relationship between the effective tracking target and the surrounding environment, combined with the temporal changes of the local deformation descriptor, logical constraint deduction is performed to obtain the expected behavior constraint.
[0027] Optionally, the control sphere first analyzes the relative topological relationships of its skeleton nodes based on the local deformation descriptor. These relative topological relationships refer to the connection methods, relative positional ratios, and spatial arrangement relationships between skeleton nodes, specifically including direct connection relationships, length ratio relationships, joint angle relationships, and topological connectivity relationships. Then, based on the relative topological relationships of the skeleton nodes, the rationality of the human body structure is verified, and valid tracking targets that conform to the characteristics of the human body structure are selected, as detailed in steps 4011 to 4015.
[0028] Optionally, based on the obtained effective tracking target, the monitoring ball analyzes the spatial relationship between the effective tracking target and the surrounding environment (such as transmission lines, towers, etc.), and then combines the temporal change characteristics of the local deformation descriptor with the video stream time series to derive the expected behavior constraints of the effective tracking target through logical constraint deduction, as in steps 4021 to 4026. The expected behavior constraints are reasonable ranges and constraints set for the subsequent possible movement behavior and posture changes of the effective tracking target based on the operation specifications of high-altitude power grid operations, environmental characteristics, human physiological structure laws, and the movement laws of the effective tracking target.
[0029] Step 50: Construct a dynamic search space based on the instantaneous motion trend and expected behavior constraints of the target being effectively tracked, and perform reverse geometric mapping and consistency verification in the dynamic search space to obtain the target tracking result.
[0030] Optionally, the control ball extracts the instantaneous motion trend of the effectively tracked target. This instantaneous motion trend refers to the immediate motion characteristics of the effectively tracked target in the current video frame, such as its motion direction, speed, and acceleration. Combined with expected behavior constraints, a dynamic search space is constructed for the position of the effectively tracked target in the next frame. This dynamic search space is an image region range that includes the possible positions of the target in the next frame, defined based on the instantaneous motion trend and behavior constraints of the target. Finally, the effectively tracked target is subjected to reverse geometric mapping within this dynamic search space, mapping the target features under a unified geometric reference back to the image coordinate system of the original video frame. The consistency of the mapped target features is then checked to obtain an accurate target tracking result, as detailed in steps 501 to 505. The target tracking result is a set of information such as the position coordinates, motion trajectory, and attitude changes of the effectively tracked target in consecutive video frames.
[0031] The embodiments of the present invention solve the problems of tracking offset and tracking loss that occur in existing methods due to their reliance on fixed global feature templates in scenarios such as complex backgrounds, small targets, and changing target features. This improves the accuracy and stability of target tracking in high-altitude power grid operations and meets the needs of high-precision safety monitoring of high-altitude power grid operations.
[0032] Optionally, the process of steps 101 to 105 includes: Step 101: Calculate the geometric centroid coordinates of all foreground pixels based on the two-dimensional coordinates of all foreground pixels in the candidate target image block, and determine the geometric centroid coordinates as the topological center reference point of the candidate target image block.
[0033] Optionally, the control sphere extracts the two-dimensional coordinates of all foreground pixels within the candidate target image patch. These two-dimensional coordinates refer to the pixel position coordinates with the top-left corner of the candidate target image patch as the origin, the horizontal direction as the abscissa, and the vertical direction as the ordinate. The abscissas of all foreground pixels are summed, and the sum is divided by the total number of foreground pixels to obtain the average abscissa of all foreground pixels. Simultaneously, the ordinates of all foreground pixels are summed, and the sum is divided by the total number of foreground pixels to obtain the average ordinate of all foreground pixels. The coordinates formed by the average abscissa and the average ordinate are the geometric centroid coordinates of all foreground pixels. These geometric centroid coordinates refer to the position center coordinates of all foreground pixels within the candidate target image patch, and the control sphere directly determines these geometric centroid coordinates as the topological center reference point of the candidate target image patch.
[0034] In one embodiment, the control ball extracts a total of 6,000 foreground pixels within a candidate target image block for high-altitude operations in a power grid. The two-dimensional coordinates of each foreground pixel are extracted one by one, yielding specific coordinates such as (21, 35), (21, 36), and (22, 35). The control ball sums the abscissas of the 6,000 foreground pixels to obtain 120,000, which is divided by 6,000 to obtain an average abscissa of 20. The control ball sums the ordinates of the 6,000 foreground pixels to obtain 180,000, which is divided by 6,000 to obtain an average ordinate of 30. Thus, the geometric centroid coordinates are (20, 30), and these coordinates are determined as the topological center reference point of the candidate target image block.
[0035] Step 102: Based on the topological center reference point, emit rays in the four boundary directions of the candidate target image block and detect the farthest intersection point of the rays with the foreground pixel point set to obtain a four-way extreme boundary point set including the farthest intersection point of the upper boundary, the farthest intersection point of the lower boundary, the farthest intersection point of the left boundary, and the farthest intersection point of the right boundary.
[0036] Optionally, the control ball uses the topological center reference point of the candidate target image patch as the emission origin and emits rays in the four boundary directions of the candidate target image patch. The four boundary directions are specifically the vertical upward direction, the vertical downward direction, the horizontal left direction, and the horizontal right direction. A ray refers to a line connecting pixels that extends infinitely in a single direction from the emission origin. Foreground pixels are detected for each ray, and the foreground pixels on the ray are identified one by one. The foreground pixel farthest from the emission origin on each ray is determined as the farthest intersection point between the ray and the set of foreground pixels in that direction. Specifically, the farthest intersection point of the vertical upward ray is the farthest intersection point of the upper boundary, the farthest intersection point of the vertical downward ray is the farthest intersection point of the lower boundary, the farthest intersection point of the horizontal left ray is the farthest intersection point of the left boundary, and the farthest intersection point of the horizontal right ray is the farthest intersection point of the right boundary. Finally, the farthest intersection points of the upper boundary, lower boundary, left boundary, and right boundary are integrated to obtain a set of four-way extreme boundary points. The set of four-way extreme boundary points refers to the set of points that can represent the farthest distribution positions of foreground pixels in four orthogonal directions within the candidate target image block.
[0037] Continuing with the embodiment of step 101, the control ball emits rays in four boundary directions—vertically upward, vertically downward, horizontally to the left, and horizontally to the right—using the topology center reference point (20, 30) as the emission origin. The detection finds that the farthest foreground pixel on the vertically upward ray is (20, 10), which is determined as the farthest intersection point on the upper boundary; the farthest foreground pixel on the vertically downward ray is (20, 50), which is determined as the farthest intersection point on the lower boundary; the farthest foreground pixel on the horizontally to the left ray is (5, 30), which is determined as the farthest intersection point on the left boundary; and the farthest foreground pixel on the horizontally to the right ray is (35, 30), which is determined as the farthest intersection point on the right boundary. The control ball integrates these four intersection points to obtain a set of four extreme boundary points: {(20, 10), (20, 50), (5, 30), (35, 30)}.
[0038] Step 103: Based on the vertical Euclidean distance between the farthest intersection point of the upper boundary and the farthest intersection point of the lower boundary in the four-way extreme boundary point set, determine the vertical topological span of the candidate target image patch, and based on the horizontal Euclidean distance between the farthest intersection point of the left boundary and the farthest intersection point of the right boundary in the four-way extreme boundary point set, determine the horizontal topological span of the candidate target image patch.
[0039] Optionally, the control ball extracts the two-dimensional coordinates of the farthest intersection point of the upper boundary and the farthest intersection point of the lower boundary within the four-way extreme boundary point set, and calculates the vertical Euclidean distance between the two intersection points. This vertical Euclidean distance refers to the straight-line distance between two points in the vertical direction in a two-dimensional coordinate system. It is calculated by taking the absolute value of the difference between the ordinates of the two points, and directly determining this vertical Euclidean distance as the vertical topological span of the candidate target image patch. That is, the vertical topological span refers to the maximum distribution span of foreground pixels within the candidate target image patch in the vertical direction. Simultaneously, the control ball extracts the two-dimensional coordinates of the farthest intersection point of the left boundary and the farthest intersection point of the right boundary within the four-way extreme boundary point set, and calculates the horizontal Euclidean distance between the two intersection points. The horizontal Euclidean distance refers to the straight-line distance between two points in the horizontal direction in a two-dimensional coordinate system. It is calculated by taking the absolute value of the difference between the abscissas of the two points, and directly determining this horizontal Euclidean distance as the horizontal topological span of the candidate target image patch. The horizontal topological span refers to the maximum distribution span of foreground pixels within the candidate target image patch in the horizontal direction.
[0040] Continuing with the embodiment of step 102, the control ball extracts the ordinates of the farthest intersection point (20, 10) on the upper boundary and the farthest intersection point (20, 50) on the lower boundary. The absolute value of the difference between the ordinates is |50-10|=40. This value is determined as the vertical topological span of the candidate target image patch, i.e., the vertical topological span is 40 pixels. Simultaneously, the abscissas of the farthest intersection point (5, 30) on the left boundary and the farthest intersection point (35, 30) on the right boundary are extracted. The absolute value of the difference between the abscissas is |35-5|=30. This value is determined as the horizontal topological span of the candidate target image patch, i.e., the horizontal topological span is 30 pixels.
[0041] Step 104: Calculate the first geometric scaling ratio in the vertical direction and the second geometric scaling ratio in the horizontal direction based on the preset fixed height and preset fixed width combined with the vertical topological span and the horizontal topological span.
[0042] Optionally, the control ball retrieves a preset fixed height and a preset fixed width. The preset fixed height refers to the pixel height of the standard-scale image block in the unified geometric reference in the vertical direction, and the preset fixed width refers to the pixel width of the standard-scale image block in the unified geometric reference in the horizontal direction. Both are fixed values preset by the control ball. Dividing the preset fixed height by the vertical topological span of the candidate target image block yields a first geometric scaling ratio in the vertical direction. This first geometric scaling ratio refers to the scaling factor required for the candidate target image block to adapt to the unified geometric reference in the vertical direction. Simultaneously, the control ball divides the preset fixed width by the horizontal topological span of the candidate target image block, yielding a second geometric scaling ratio in the horizontal direction. This second geometric scaling ratio refers to the scaling factor required for the candidate target image block to adapt to the unified geometric reference in the horizontal direction.
[0043] Continuing with the embodiment of step 103, the control ball retrieves the preset fixed height of 100 pixels and the preset fixed width of 75 pixels. The preset fixed height of 100 pixels is divided by the vertical topological span of 40 pixels to calculate 100÷40=2.5, which is the first geometric scaling ratio in the vertical direction. The preset fixed width of 75 pixels is divided by the horizontal topological span of 30 pixels to calculate 75÷30=2.5, which is the second geometric scaling ratio in the horizontal direction.
[0044] Step 105: Based on the topological center reference point of the candidate target image patch, perform topological mapping by combining the first geometric scaling ratio and the second geometric scaling ratio to obtain a standard-scale image patch.
[0045] Optionally, the control ball uses the topological center reference point of the candidate target image block as the core reference, and combines the first geometric scaling ratio in the vertical direction and the second geometric scaling ratio in the horizontal direction to perform topological mapping transformation on all pixels in the candidate target image block, and finally obtains a standard-scale image block that meets the requirements of a unified geometric reference, as in steps 1051 to 1054.
[0046] In this embodiment of the invention, the geometric centroid of the foreground pixel is used to determine the topological center reference point. The actual topological span of the target is determined by detecting the four-way extreme boundary points. Combined with a preset fixed size, a precise scaling ratio is calculated to unify various candidate target image blocks under a fixed geometric reference. This eliminates the geometric scale differences between different candidate target image blocks and solves the problem of inconsistent image feature scale caused by the low proportion of target pixels and obvious small target features in high-altitude power grid operation scenarios.
[0047] Optionally, the process of steps 1051 to 1054 includes: Step 1051: Based on the topological center reference point of the candidate target image patch, construct a standard topological mapping grid covering the entire region of the candidate target image patch using a first geometric scaling ratio and a second geometric scaling ratio.
[0048] Optionally, the control sphere uses the topological center reference point of the candidate target image block as the grid center reference point, calls the first geometric scaling ratio in the vertical direction and the second geometric scaling ratio in the horizontal direction, and constructs a standard topological mapping grid covering the entire area of the candidate target image block according to the pixel size of the standard-scale image block corresponding to the unified geometric reference. This standard topological mapping grid refers to a pixel grid structure centered on the topological center reference point and divided according to a preset scaling ratio, corresponding one-to-one with the pixel units of the standard-scale image block. The grid unit is the smallest square pixel area constituting the standard topological mapping grid, and the pixel side length of each grid unit is a fixed value. Covering the entire area of the candidate target image block means that the boundary range of the grid completely includes all pixels of the candidate target image block, and no pixels exceed the grid boundary. During the construction process, with the topological center reference point as the origin, the grid is divided unit by unit along the horizontal left-right and vertical up-down directions according to the first and second geometric scaling ratios, until the grid boundary completely covers the area enclosed by the four extreme boundary points of the candidate target image block.
[0049] Continuing with the embodiment of step 104, the control ball uses the topology center reference point (20, 30) as the grid center reference point, calls the first geometric scaling ratio of 2.5 and the second geometric scaling ratio of 2.5, and constructs a standard topology mapping grid according to the size of a standard scale image block of 500×500 pixels. The pixel side length of the grid unit is 1 pixel. With (20, 30) as the origin, the grid is divided pixel by pixel along the horizontal left and right directions with a scaling ratio of 2.5, and along the vertical up and down directions with a scaling ratio of 2.5, until the grid boundary completely covers the entire area enclosed by the four extreme boundary points (20, 10), (20, 50), (5, 30), and (35, 30) of the candidate target image block, and finally obtains a standard topology mapping grid that covers the entire area of the candidate target image block.
[0050] Step 1052: Based on the position of each foreground pixel in the candidate target image block in the grid cell of the standard topological mapping grid, establish a direct coordinate mapping table of the foreground pixels in the candidate target image block.
[0051] Optionally, the control sphere retrieves the original two-dimensional coordinates of each foreground pixel in the candidate target image block, and determines the grid cell position of each foreground pixel in the standard topological mapping grid. The grid cell position refers to the specific grid cell number and corresponding coordinate range in the standard topological mapping grid where the original two-dimensional coordinates of the foreground pixel fall. The control sphere assigns a unique grid cell number to each grid cell of the standard topological mapping grid, and this number corresponds one-to-one with the pixel position number of the standard-scale image block. Using "original two-dimensional coordinates of foreground pixels in the candidate target image block - grid cell position - target pixel position in the standard-scale image block" as the mapping relationship, the above correspondence of all foreground pixels is integrated to establish a structured direct coordinate mapping relationship table for foreground pixels in the candidate target image block. This direct coordinate mapping relationship table refers to a structured data table that records the one-to-one correspondence between foreground pixels from their original image coordinates to their standard-scale image coordinates.
[0052] Continuing with the embodiment of step 1051, the control ball extracts the original two-dimensional coordinates of all foreground pixels (20, 10), (20, 30), (35, 30), etc., within the candidate target image block. It then determines the position of each point within its corresponding grid cell in the standard topological mapping grid. For example, foreground pixel (20, 10) falls into the region with grid cell number G1001, which corresponds to the target pixel position (250, 100) of the standard-scale image block; foreground pixel (20, 30) falls into the region with grid cell number G2500, which corresponds to the target pixel position (250, 250) of the standard-scale image block. The control ball integrates the "original two-dimensional coordinates - grid cell position - target pixel position corresponding to the standard-scale image block" information of all foreground pixels to establish a direct coordinate mapping relationship table such as (20, 10) - G1001 - (250, 100), (20, 30) - G2500 - (250, 250).
[0053] Step 1053: Based on the direct coordinate mapping table, fill the grayscale or color values of the foreground pixels in each grid cell of the candidate target image block to the corresponding target pixel positions in the standard scale image to obtain the initial scale image block for foreground texture information topological transfer.
[0054] Optionally, the control ball retrieves the direct coordinate mapping table to extract the grayscale or color values of all foreground pixels in each grid cell of the candidate target image block. The grayscale value refers to the brightness value of a pixel in a single-channel image, ranging from 0 to 255, where 0 represents pure black and 255 represents pure white. The color value refers to the color value of a pixel in a color image, represented by red, green, and blue channels, with each channel ranging from 0 to 255. According to the correspondence in the direct coordinate mapping table, the grayscale or color values of the foreground pixels in each grid cell are precisely filled to the corresponding target pixel positions in the standard-scale image. The filling method is a one-to-one pixel value assignment, that is, the grayscale or color value of a single foreground pixel in the candidate target image block is directly assigned to the corresponding single target pixel in the standard-scale image. After completing the filling of the grayscale or color values of all foreground pixels, an initial-scale image block with foreground texture information topological transfer is obtained. This initial-scale image block refers to the prototype of a standard-scale image block that has only completed the foreground texture information filling and has not filled the background pixels, containing all the foreground texture features of the candidate target image block.
[0055] Continuing with the embodiment of step 1052, the control ball retrieves the direct coordinate mapping table, extracts the gray value 180 of the foreground pixel (20, 10) in grid cell G1001 of the candidate target image block, and fills the gray value 180 into the target pixel position (250, 100) in the standard scale image according to the corresponding relationship; extracts the gray value 200 of the foreground pixel (20, 30) in grid cell G2500, and fills the gray value 200 into the target pixel position (250, 250) in the standard scale image, and so on, filling all the gray values of the foreground pixels according to the mapping relationship, and finally obtains the initial scale image block containing only foreground texture information.
[0056] Step 1054: Based on the initial scale image block, fill the remaining blank pixel positions without foreground texture information with default background values to obtain a standard scale image block.
[0057] Optionally, the control sphere identifies all remaining blank pixel locations in the initial-scale image patch that have not been filled with foreground texture information. These remaining blank pixel locations refer to pixel locations in the initial-scale image patch that have not been assigned grayscale or color values by foreground pixels; these locations have no texture information. A preset background default value is then invoked. This default value refers to a grayscale or color value pre-set by the control sphere to fill the blank pixel locations; it is a fixed value that conforms to the pixel feature adaptation requirements of the power grid high-altitude operation scenario background. The background default value is then uniformly filled into all remaining blank pixel locations. The filling method is to uniformly assign the same background default value to all blank pixel locations. After filling all pixel locations, a standard-scale image patch conforming to the unified geometric reference requirements is obtained.
[0058] Continuing with the embodiment of step 1053, the control ball identifies all unfilled blank pixel positions in the initial scale image block, except for the pixel positions (250, 100), (250, 250), which have been filled with foreground gray values. It then calls the preset background default gray value 50 and uniformly fills all blank pixel positions with the gray value 50. After filling all pixel positions, a standard scale image block of 500×500 pixels that conforms to a unified geometric benchmark is obtained.
[0059] This invention, through the construction of a standard topology mapping grid corresponding one-to-one with standard-scale image blocks and the establishment of a direct coordinate mapping table, solves the problem of incomplete mapping caused by the scattered distribution of foreground pixels in small target scenarios of high-altitude power grid operations. It also eliminates the interference of background pixel differences between different candidate target image blocks on subsequent feature extraction, thereby achieving lossless topology migration from candidate target image blocks to standard-scale image blocks.
[0060] Optionally, the process of steps 4011 to 4015 includes: Step 4011: Extract directly connected skeleton node pairs based on the coordinates of all skeleton nodes and their node type identifiers in the local deformation descriptor to obtain the initial limb connection group.
[0061] Optionally, the control ball retrieves the coordinates of all skeleton nodes contained in the local deformation descriptor, and simultaneously extracts the node type identifier corresponding to each skeleton node. This node type identifier refers to the feature identifier used to distinguish the human body part to which the skeleton node belongs. This identifier corresponds one-to-one with the preset human body parts, specifically including identifiers such as head vertex, left shoulder endpoint, right shoulder endpoint, left elbow endpoint, right elbow endpoint, left wrist endpoint, right wrist endpoint, hip midpoint, left knee endpoint, right knee endpoint, left ankle endpoint, and right ankle endpoint. Based on the limb connection rules of normal human physiological structure, the skeleton nodes with node type identifiers are matched to extract all directly connected skeleton node pairs. A directly connected skeleton node pair refers to two skeleton nodes that conform to human physiological structure and are directly connected in the limb. Each pair of directly connected skeleton nodes constitutes a limb connection. All limb connections are integrated to form an initial limb connection group.
[0062] In one embodiment, the control ball retrieves a local deformation descriptor containing 12 skeletal nodes, extracts the coordinates of each node and its corresponding node type identifier, including the head vertex, left shoulder endpoint, right shoulder endpoint, left elbow endpoint, right elbow endpoint, left wrist endpoint, right wrist endpoint, hip midpoint, left knee endpoint, right knee endpoint, left ankle endpoint, and right ankle endpoint. Based on the limb connection rules of normal human physiological structure, nine directly connected skeletal node pairs are extracted: head vertex-hip midpoint, left shoulder endpoint-left elbow endpoint, left elbow endpoint-left wrist endpoint, right shoulder endpoint-right elbow endpoint, right elbow endpoint-right wrist endpoint, hip midpoint-left knee endpoint, left knee endpoint-left ankle endpoint, hip midpoint-right knee endpoint, and right knee endpoint-right ankle endpoint. Each node pair constitutes a limb connection. These nine limb connections are integrated to obtain the initial limb connection group.
[0063] Step 4012: Based on the coordinates of the two endpoint skeleton nodes of each limb connection in the initial limb connection group, calculate the set of length ratios between the length of each limb connection and the reference length, and compare each set of length ratios with the preset reasonable proportion range of human anatomy to select qualified limb connection groups that pass the length ratio verification.
[0064] Optionally, the control ball retrieves the coordinates of the two endpoint skeleton nodes of each limb connection within the initial limb connection group. Using the Euclidean distance calculation method in a two-dimensional coordinate system, it calculates the actual length of each limb connection, which refers to the straight-line distance between the two endpoint skeleton nodes. It also retrieves a preset reference length, a fixed value representing the baseline limb length for comparison, set according to human anatomy standards. The actual length of each limb connection is divided by the reference length to obtain the length ratio for each connection, and all length ratios are integrated to form a length ratio set. Simultaneously, a preset reasonable anatomical proportion range is retrieved. This range refers to a numerical range that conforms to the normal limb length proportions, set based on human anatomical characteristics. Different types of limb connections correspond to specific reasonable proportion ranges. Each length ratio in the length ratio set is compared one by one with the reasonable anatomical proportion range for its corresponding limb connection. Unqualified limb connections with length ratios exceeding the reasonable proportion range are eliminated, while those with length ratios within the reasonable proportion range are retained. All retained limb connections are integrated to form a qualified limb connection group that passes the length ratio verification.
[0065] Continuing with the embodiment of step 4011, the control ball extracts the coordinates of the two endpoints of the left shoulder endpoint-left elbow endpoint limb connection in the initial limb connection group, calculates the actual length of the limb connection as 30 pixels, retrieves the reference length as 20 pixels, and calculates the length ratio as 1.5; extracts the coordinates of the two endpoints of the left elbow endpoint-left wrist endpoint limb connection, calculates the actual length as 28 pixels, and calculates the length ratio as 1.4, and so on to obtain the length ratios of all 9 limb connections, forming a length ratio set {1.2, 1.5, 1.4, 1.5, 1.3, 1.6, 1.5, 1.6, 1.4}. Retrieve the preset reasonable proportion range of human anatomy, where the reasonable proportion range for arm limbs is 1.2 to 1.6 and the reasonable proportion range for leg limbs is 1.4 to 1.7. Compare each length ratio with the corresponding range one by one. Since the length ratio of all limb connections is within the corresponding reasonable proportion range, all 9 limb connections in the initial limb connection group are retained to form a qualified limb connection group.
[0066] Step 4013: Based on the coordinates of three consecutive connected skeleton nodes in the qualified limb connection group, construct the joint angle geometry structure composed of two adjacent limbs. Then, compare the internal angle value of each joint angle geometry structure with the preset human joint physiological activity angle threshold one by one to select qualified joint structure groups that pass the angle verification.
[0067] Optionally, the control ball extracts all three consecutively connected skeleton nodes based on all limb connections within a qualified limb connection group and the joint connection characteristics of human physiological structure. These three consecutively connected skeleton nodes refer to three skeleton nodes where the first and second nodes form a limb connection, and the second and third nodes form an adjacent limb connection, with the second node being the joint core node. Based on the two-dimensional coordinates of these three consecutively connected skeleton nodes, the control ball constructs a joint angle geometry structure composed of two adjacent limb connections. This joint angle geometry structure refers to the geometric angle structure formed with the joint core node as the vertex and the two adjacent limb connections as the edges. The control ball also calculates the internal angle value of each joint angle geometry structure. The internal angle value refers to the inner angle formed by the two adjacent limb connections at the joint core node, with a value range of 0 to 180 degrees. Simultaneously, the control ball retrieves a preset human joint physiological activity angle threshold. This threshold refers to an angle value range set according to the characteristics of human joint physiological activity that conforms to the normal range of human joint activity. Different types of human joints correspond to specific physiological activity angle thresholds. The internal angle value of each joint's geometric structure is compared one by one with the corresponding human joint physiological activity angle threshold. Unqualified joint structures with internal angle values exceeding the threshold range are eliminated, while joint structures with internal angle values within the threshold range are retained. All retained joint structures are integrated to form a qualified joint structure group that has passed the angle verification.
[0068] Continuing with the embodiment of step 4012, the control ball extracts four groups of three consecutive connected skeletal nodes from the qualified limb connection group: left shoulder endpoint-left elbow endpoint-left wrist endpoint, right shoulder endpoint-right elbow endpoint-right wrist endpoint, hip midpoint-left knee endpoint-left ankle endpoint, and hip midpoint-right knee endpoint-right ankle endpoint. Using the core node of each group as the vertex, four joint angle geometric structures are constructed: the left elbow joint, right elbow joint, left knee joint, and right knee joint angle structures. The calculated internal angle values are 120 degrees, 115 degrees, 130 degrees, and 125 degrees, respectively. Preset human joint physiological activity angle thresholds are retrieved, with the elbow joint physiological activity angle threshold being 90 to 150 degrees and the knee joint physiological activity angle threshold being 100 to 160 degrees. Each internal angle value is compared with its corresponding threshold. It is found that all four angle values are within the corresponding threshold range; therefore, all four joint structures are retained, forming a qualified joint structure group.
[0069] Step 4014: Based on all skeleton nodes in the qualified joint structure group, reconstruct the connection paths between nodes, remove isolated skeleton nodes that are disconnected or floating limb branches that are not connected to the main node, and obtain a connected skeleton node group with complete topological connectivity.
[0070] Optionally, the control ball reconstructs the connectivity paths between nodes based on all skeletal nodes within the qualified joint structure group and their connection relationships within the group. These connectivity paths conform to human physiological structure and connect all skeletal nodes into a unified whole. The reconstructed connectivity paths are then subjected to integrity checks, identifying and removing isolated skeletal nodes that are broken, as well as suspended limb branches not connected to the main skeletal nodes. Isolated skeletal nodes are those not connected to any other skeletal nodes and existing in a separate state; suspended limb branches are those not connected to the main skeletal nodes and existing independently. Main skeletal nodes are key nodes constituting the core skeletal structure of the human body, including the head apex, hip midpoint, left and right shoulder endpoints, and left and right knee endpoints. After removing isolated skeletal nodes and suspended limb branches, a connected skeletal node group with complete topological connectivity is obtained, consisting of interconnected skeletal nodes.
[0071] Continuing with the embodiment of step 4013, the control ball extracts 12 skeleton nodes from the qualified joint structure group. Based on the joint connection relationships, the connectivity paths between the nodes are reconstructed, forming a complete connectivity path with the head apex and hip midpoint as the main trunk, extending to the left and right arms and legs. An integrity check of this connectivity path reveals no broken, isolated skeleton nodes or suspended limb branches not connected to the main trunk node. Therefore, these 12 interconnected skeleton nodes are directly integrated to obtain a connected skeleton node group with complete topological connectivity.
[0072] Step 4015: Based on the connected skeleton node group, the original local deformation descriptor is retrieved in reverse, mismatched node data and connection relationships are located and removed to obtain the remaining node data and connection relationships. The remaining node data and connection relationships are then re-encapsulated to obtain the effective tracking target.
[0073] Optionally, the control ball uses the connected skeleton node group as the retrieval basis to reverse-search the original local deformation descriptor constructed in step 30, locating node data and connections in the original local deformation descriptor that do not match the connected skeleton node group. This mismatched node data and connections refer to the coordinates and node type identifiers of skeleton nodes not included in the connected skeleton node group in the original local deformation descriptor, as well as the corresponding limb connections. The located mismatched node data and connections are removed, retaining the remaining node data and connections in the original local deformation descriptor that match the connected skeleton node group. This remaining node data and connections include the coordinates and node type identifiers of all skeleton nodes within the connected skeleton node group, as well as the corresponding qualified limb connections and qualified joint structures. Finally, according to the encapsulation specifications of the local deformation descriptor, the remaining node data and connections are re-encapsulated to form feature description data that meets the requirements of human structural rationality. This feature description data is the effective tracking target.
[0074] Continuing with the embodiment of step 4014, the control ball, based on the connected skeleton node group containing 12 skeleton nodes, reversely retrieves the original local deformation descriptor. The detection shows that all node data and connections in the original local deformation descriptor match the connected skeleton node group, and there is no mismatched data that needs to be removed. Following the encapsulation specifications of the local deformation descriptor, the control ball re-encapsulates the node data and connections in the original local deformation descriptor to form feature description data that meets the requirements of human body structure rationality. This data is the effective tracking target.
[0075] This invention verifies the rationality of human body structure from three dimensions: limb connection, joint angle, and topological connectivity. It effectively eliminates invalid nodes and erroneous connections in the local deformation descriptor caused by factors such as complex backgrounds, small targets, and lighting fluctuations in high-altitude power grid operations. This avoids misidentifying background interference pixels as human skeleton nodes and misjudging deformed structures as valid targets, and solves the problem of target misjudgment that is prone to occur in existing methods due to global feature template interference.
[0076] Optionally, the process of steps 4021 to 4026 includes: Step 4021: Extract the relative displacement vector direction of the skeleton node coordinates in consecutive frames based on the skeleton node coordinate sequence in the local deformation descriptor to obtain the limb motion tendency vector group, and identify the main axis direction of limb extension based on the vector direction with the largest magnitude in the limb motion tendency vector group to obtain the locked potential action force point.
[0077] Optionally, the control ball extracts the coordinate sequence of corresponding skeleton nodes within consecutive video frames based on local deformation descriptors. For each pair of directly connected skeleton nodes, it calculates the coordinate difference between the node pair in adjacent frames to obtain the relative displacement vector of the limb node. All relative displacement vectors of the limb nodes are then integrated to form a limb motion tendency vector group. Here, the skeleton node coordinate sequence refers to the set of two-dimensional pixel coordinates corresponding to the same skeleton node in multiple consecutive frames of images for the effectively tracked target; the relative displacement vector refers to the vector reflecting the direction and magnitude of the position change of the limb node in adjacent frames, including two core elements: displacement direction and displacement magnitude.
[0078] Optionally, the control ball traverses all vectors in the limb movement tendency vector group and identifies the vector direction with the largest magnitude. Magnitude refers to the displacement amplitude relative to the displacement vector. This vector direction is the most important limb movement extension direction of the effectively tracked target. Based on this main movement extension direction, the starting core node of the limb movement is located in reverse, and this core node is identified as the locked potential action force point. This potential action force point refers to the core node that is the source of power for the current limb movement trend of the effectively tracked target.
[0079] In one embodiment, the control ball retrieves the local deformation descriptors of the effectively tracked target from three consecutive frames of images, and extracts the coordinate sequences of the left shoulder endpoint, left elbow endpoint, and left wrist endpoint. The coordinate difference between the left shoulder endpoint and the left elbow endpoint in adjacent frames is calculated, yielding the relative displacement vector of the left arm node as (2, 3), and the coordinate difference between the left elbow endpoint and the left wrist endpoint as (1, 2). The corresponding relative displacement vectors for the right half of the limbs are (3, 4) and (2, 3), and the relative displacement vectors for the torso and legs are (1, 1), (2, 2), and (1, 2). All vectors are integrated to form a limb movement tendency vector group, and the vector with the largest modulus is identified as (3, 4) corresponding to the right shoulder endpoint - right elbow endpoint, with its direction being diagonally upward to the right. This determines the right shoulder endpoint as the potential point of force for the locked action.
[0080] Step 4022: Based on the geometric center position calculated by combining the trunk center node and bilateral hip joint nodes of the effective tracking target with the potential action force point, an equivalent balance reference point is obtained. Then, based on the coordinates of the pole structure surface in the surrounding environmental topology elements and the coordinates of the end skeleton node of the effective tracking target, a spatial proximity projection is performed to obtain a candidate environmental contact point group.
[0081] Optionally, the control ball extracts the coordinates of the trunk center node and bilateral hip joint nodes from the effectively tracked target. The trunk center node refers to the skeletal node constituting the core of the target's trunk, specifically the midpoint of the hip. The bilateral hip joint nodes refer to the key skeletal nodes connecting the trunk and legs, namely the left hip node and the right hip node. Based on the coordinates of the trunk center node and bilateral hip joint nodes, combined with the coordinates of potential force points, the spatial center coordinates are calculated using equal or preset weights, and this geometric center position is determined as the equivalent balance reference point. This equivalent balance reference point is a virtual reference point representing the overall posture balance state of the effectively tracked target.
[0082] Optionally, the monitoring sphere retrieves topological elements of the surrounding environment, specifically the geometric features of fixed structures such as poles, towers, crossarms, and ladders in the power grid high-altitude operation scenario. This includes the surface coordinates of the pole structure, which refer to the set of pixel coordinates constituting the surface of the pole entity. Simultaneously, it extracts the coordinates of the end-effector skeleton nodes of the effective tracking target. These end-effector skeleton nodes refer to key skeleton nodes at the limb ends, specifically the left wrist endpoint, right wrist endpoint, left ankle endpoint, and right ankle endpoint. Each end-effector skeleton node coordinate is then spatially proximately projected onto the pole structure surface. This spatial proximate projection maps the end-effector node coordinates to the nearest corresponding coordinate point on the pole structure surface, maintaining the spatial topological relationships unchanged during the projection process. All projected coordinate points are then integrated to form a candidate environmental contact point group.
[0083] Continuing with the embodiment of step 4021, the control ball extracts the coordinates of the midpoint of the hip (250, 200), the left hip node (230, 220), and the right hip node (270, 220) of the effective tracking target. Combined with the coordinates of the right shoulder endpoint (280, 150) of the potential action force point, the geometric center position of the three is calculated to be (250, 198), which is determined as the equivalent balance reference point. The control ball retrieves the coordinates of the surrounding tower structure surface and extracts the left wrist endpoint (290, 180), right wrist endpoint (310, 180), left ankle endpoint (240, 300), and right ankle endpoint (260, 300) of the effective tracking target. The coordinates of each end node are projected onto the tower structure surface to obtain the projected coordinate points (290, 182), (310, 181), (240, 302), and (260, 301), respectively. These points are integrated to form a candidate environmental contact point group {(290, 182), (310, 181), (240, 302), (260, 301)}.
[0084] Step 4023: Construct a dynamic support topology based on the minimum convex hull boundary formed by the candidate environmental contact point group, determine whether the equivalent equilibrium reference point falls inside the dynamic support topology domain, and obtain the attitude stability state of the target.
[0085] Optionally, the control ball sorts all contact points in the candidate environmental contact point group and connects them sequentially in a clockwise direction to form a closed convex polygon. This convex polygon is the minimum convex hull boundary, which refers to the boundary of the smallest convex polygon that can completely enclose the candidate environmental contact point group. This convex polygon boundary and the candidate environmental contact point group together constitute the dynamic support topology. The dynamic support topology refers to the spatial topology that characterizes the contact support relationship between the effective tracking target and the tower structure, and the area it encloses is the dynamic support topology domain. Then, the control ball determines whether the equivalent equilibrium reference point falls within the dynamic support topology domain. The determination rule is: if the coordinates of the equivalent equilibrium reference point are within the area enclosed by the minimum convex hull boundary, the attitude of the effective tracking target is determined to be in a stable state; if the coordinates of the equivalent equilibrium reference point are outside the area enclosed by the minimum convex hull boundary, the attitude of the effective tracking target is determined to be in an unstable state. Based on this determination result, the attitude stability state of the target is obtained, i.e., the attitude stability state includes both stable and unstable states.
[0086] Continuing with the embodiment of step 4022, the control ball constructs a minimum convex hull boundary based on the candidate environmental contact point group {(290, 182), (310, 181), (240, 302), (260, 301)}, forming a convex quadrilateral connected by these four points, constituting a dynamic support topology. The coordinates of the equivalent equilibrium reference point (250, 198) are determined, and it is found to be within the area enclosed by the convex quadrilateral. Therefore, the attitude stability state of the effectively tracked target is determined to be an attitude stable state.
[0087] Step 4024: Based on the attitude stability state and the normal direction of the tower structure surface in the surrounding environment topology, calculate the projection components of the equivalent balance reference point and the potential action force point on the normal direction of the tower structure surface to obtain the attitude instability trend vector. Based on the extension direction of the attitude instability trend vector, filter the regions where the surface normal direction and the attitude instability trend vector have an inverse blocking topological relationship to obtain environmental contact candidate regions.
[0088] Optionally, based on the attitude stability state, the control ball retrieves the normal direction of the tower structure surface from the surrounding environmental topology elements. This normal direction refers to the vertical direction of the tower structure surface at the corresponding coordinate point, and is fixed spatial direction information. The projection components of the equivalent equilibrium reference point and the potential action point on the normal direction of the tower structure surface are calculated. These projection components are scalar values obtained by projecting the coordinate vectors of the equivalent equilibrium reference point and the potential action point onto the normal direction of the tower structure surface. By calculating the difference between the two projection components, the attitude instability trend vector is obtained. This vector characterizes the trend and magnitude of the effective tracking target's attitude developing towards instability. If the attitude stability state is stable, the magnitude of the attitude instability trend vector is small, and its direction points inside the dynamic support topology domain; if the attitude stability state is unstable, the magnitude of the attitude instability trend vector is large, and its direction points outside the dynamic support topology domain.
[0089] Optionally, the control ball analyzes the extension direction of the attitude instability trend vector and filters out regions where the normal direction of the tower structure surface has an inverse blocking topology relationship with the extension direction. The inverse blocking topology relationship refers to regions where the normal direction of the tower structure surface is opposite to the extension direction of the attitude instability trend vector, and can block the instability trend. The filtered regions are identified as environmental contact candidate regions, which refer to regions where the effective tracking target is to maintain attitude balance and may have actual contact with the tower structure.
[0090] Continuing with the embodiment of step 4024, the control ball retrieves the attitude stability state (such as the attitude instability state), and simultaneously retrieves the normal direction of the tower structure surface as vertically upward. The projection component of the equivalent equilibrium reference point (250, 198) in the normal direction is calculated to be 198, and the projection component of the right shoulder endpoint (280, 150) of the potential action force point is 150. The difference between the two is 48, indicating that the extension direction of the attitude instability trend vector is vertically downward. Regions where the normal direction (vertically upward) of the tower structure surface and this extension direction (vertically downward) have an inverse blocking topology relationship are selected, i.e., the vertically upward support region of the tower structure surface, and these are identified as candidate environmental contact regions.
[0091] Step 4025: Based on the coordinates of the end skeleton nodes in the environmental contact candidate region and the local deformation descriptor, spatial proximity matching is performed to obtain the actual contact feature node pairs between the target and the environment, and the physical anchoring point of the target in the environmental topology is determined based on the actual contact feature node pairs.
[0092] Optionally, the monitoring ball extracts the coordinates of all pole and tower surface coordinates within the environmental contact candidate area. It then performs spatial proximity matching between the coordinates of the end skeleton node of the effective tracking target and the pole and tower surface coordinates within the environmental contact candidate area to obtain a matching result. This spatial proximity matching refers to finding the closest corresponding point between the end skeleton node coordinates and the pole and tower surface coordinates, with the matching rule being the minimum Euclidean distance between the two points. Based on the matching result, the actual contact correspondence between the end skeleton node and the pole and tower surface is determined, forming actual contact feature node pairs. These actual contact feature node pairs refer to the coordinate pairing of the actual contact point between the end skeleton node of the effective tracking target and the pole and tower surface.
[0093] Optionally, the control ball determines the physical anchoring point of the target in the environmental topology based on the actual contact feature node pairs. The physical anchoring point refers to the key coordinate point where the target and the tower structure make actual physical contact and form a fixed support relationship. It is the core support point for the target to maintain its position stability in the high-altitude operation environment of the power grid.
[0094] Continuing with the above embodiment, the control ball determines the candidate environmental contact area as the vertically upward support area of the tower structure surface, and extracts the coordinate points of the tower structure surface within this area {(290, 180), (310, 180), (240, 300), (260, 300)}. The left wrist endpoint (290, 180), right wrist endpoint (310, 180), left ankle endpoint (240, 300), and right ankle endpoint (260, 300) of the end skeleton nodes are matched with the coordinate points of this area. It is found that the Euclidean distance is 0 for all of them, forming actual contact feature node pairs {(left wrist endpoint, 290, 180), (right wrist endpoint, 310, 180), (left ankle endpoint, 240, 300), (right ankle endpoint, 260, 300)}. The control ball determines these actual contact points as the physical anchoring points of the effective tracking target in the environmental topology.
[0095] Step 4026: Based on the physical anchor points and the limb movement tendency vector group, perform geometric constraint deduction to obtain the expected behavior constraints.
[0096] Optionally, the control ball uses physical anchor points as the core support constraint and combines the limb movement tendency vector group to perform geometric constraint deduction to obtain the expected behavior constraint, as in steps 40261 to 40264.
[0097] The embodiments of the present invention achieve accurate prediction and constraint of the movement behavior of the target, effectively avoiding the tracking deviation and tracking loss problems caused by the failure to consider the interaction between the target and the environment and insufficient prediction of the movement trend in the existing methods. It improves the logical constraint system of target tracking, ensures that the target tracking process can be dynamically adjusted according to the target movement trend, and improves the accuracy and stability of target tracking in high-altitude power grid operations.
[0098] Optionally, the processes of steps 40261 to 40264 include: Step 40261: Based on the physical anchor point and the limb motion tendency vector group, the motion trajectory of the remaining degrees of freedom after being constrained by the physical anchor point is derived to obtain the restricted motion path envelope. Based on the restricted motion path envelope, invalid motion components that exceed the physical boundary of the tower structure are eliminated to obtain the effective motion path envelope.
[0099] Optionally, the control ball retrieves the physical anchor points and the limb motion tendency vector group. Based on the fixed support constraint characteristics of the physical anchor points, the motion degree of freedom of each vector in the limb motion tendency vector group is limited. The motion trajectory corresponding to the remaining degree of freedom after constraint is derived. All the motion trajectories of the remaining degree of freedom are integrated to form a closed path range. This range is the restricted motion path envelope. That is, the restricted motion path envelope refers to the set of trajectory ranges that can be effectively tracked by the target limb after being supported and constrained by the physical anchor points.
[0100] Optionally, the monitoring sphere retrieves the physical boundary of the tower structure in the high-altitude power grid operation scenario. This physical boundary refers to the pixel coordinate boundary range of the tower's physical structure in the image coordinate system, which is the physical constraint boundary for limb movement. Motion trajectory components that exceed this physical boundary of the tower structure within the restricted motion path envelope are determined as invalid motion components. These invalid motion components refer to the portion of the limb movement trajectory that is blocked by the physical structure and cannot actually occur. These invalid motion components are then removed. The remaining motion trajectory range within the physical boundary of the tower structure after removal is the valid motion path envelope.
[0101] Continuing with the above embodiment, the control ball retrieves the effective physical anchor point group {(290, 180), (310, 180), (240, 300), (260, 300)} and the limb movement tendency vector group {(2, 3), (1, 2), (3, 4), (2, 3), (1, 1), (2, 2), (1, 2)}, derives the motion trajectory of the remaining degrees of freedom after being constrained by the anchor points, and integrates them to form a restricted motion path envelope containing the left, right, and up directions. The control ball retrieves the pixel range of the tower structure's physical boundary, which is the horizontal coordinate from 50 to 400 and the vertical coordinate from 80 to 350. It detects that the motion trajectory component in the upward direction that exceeds the vertical coordinate 80 in the restricted motion path envelope is an invalid motion component, which is then removed, finally obtaining the effective motion path envelope within the tower's physical boundary.
[0102] Step 40262: Based on the effective motion path envelope and the preset safe operation area topology range, a spatial inclusion comparison is performed to obtain the path compliance identifier. Based on the path compliance identifier, the critical motion segment that is about to cross the safety boundary is marked to obtain the critical marked motion segment.
[0103] Optionally, the monitoring ball retrieves the effective motion path envelope and the preset safe working area topology. This safe working area topology refers to the safe pixel coordinate range where workers can move their limbs, as set according to the power grid high-altitude operation safety regulations. It is the safety constraint boundary for limb movement. The effective motion path envelope and the safe working area topology are spatially compared. Point by point, it is determined whether the trajectory points of the effective motion path envelope are within the safe working area topology. Based on the comparison results, a path compliance mark is assigned to the effective motion path envelope. The path compliance mark is an indicator that represents whether the effective motion path envelope conforms to the safe working area constraints. It includes two types: compliance mark and critical mark. A compliance mark indicates that the trajectory point is completely within the safe area, while a critical mark indicates that the trajectory point is close to or about to cross the safety boundary.
[0104] Optionally, the monitoring ball identifies the motion trajectory segments that are about to cross the boundary of the safe working area for each segment of the valid motion path envelope with critical markers. These trajectory segments are critical motion segments, which are then marked. The marked trajectory segments are critical marked motion segments.
[0105] Continuing with the above embodiment, the monitoring ball retrieves the effective motion path envelope and the preset safe operating area topology (horizontal coordinate 60 to 390, vertical coordinate 90 to 340). It then performs a spatial inclusion comparison between the trajectory points of the effective motion path envelope and the safe area range. It finds that some trajectory points in the envelope have horizontal coordinates close to 60 and 390 and vertical coordinates close to 90 and 340, assigning them critical markers. The monitoring ball further identifies motion trajectory segments with horizontal coordinates of 58 to 62 and 388 to 392, and vertical coordinates of 88 to 92 and 338 to 342 as critical motion segments about to cross the safety boundary. These segments are then marked to obtain the critical marked motion segments.
[0106] Step 40263: Based on the critical marker motion segment and physical anchor point, construct a rigid motion trajectory ring with the physical anchor point as the rotation center and the preset anatomical length of the corresponding limb segment as the radius. Then, based on the preset physiological activity sector of the human joint, geometrically intersect the rigid motion trajectory ring with the preset physiological activity sector to obtain the limb configuration preservation feasible region.
[0107] Optionally, the control ball retrieves the critical marker motion segments and physical anchor points. Using each physical anchor point as the rotation center, it simultaneously retrieves the preset anatomical length of the corresponding limb segment, using this length as the rotation radius. This preset anatomical length refers to the standard pixel length of different limb segments set according to human anatomy standards, and is a fixed value. A circular trajectory range around the physical anchor points is constructed; this range is the rigid motion trajectory ring. Next, preset physiological activity sectors of human joints are retrieved. These preset physiological activity sectors refer to the range of angles that a corresponding joint can move, set according to the physiological activity characteristics of human joints. Different types of human joints are matched with specific physiological activity sectors. The rigid motion trajectory ring and the corresponding preset physiological activity sectors are geometrically intersected. Geometric intersection calculation refers to the spatial calculation process of finding the overlapping area of two geometric shapes; the calculated overlapping area is the feasible region for maintaining the limb configuration.
[0108] Continuing with the above embodiment, the control ball retrieves the critical marker motion segment and the physical anchor point (310, 180). Using this physical anchor point as the rotation center, it retrieves the preset anatomical length of the corresponding right forearm segment, which is 30 pixels. A rigid motion trajectory loop is constructed with a radius of 30 pixels. The control ball retrieves the preset physiological activity sector of the human elbow joint, which is an angle sector of 90 to 150 degrees. The geometric intersection calculation is performed between the rigid motion trajectory loop and this physiological activity sector of 90 to 150 degrees to obtain the overlapping area. This area is the feasible region for maintaining the limb configuration of the right forearm.
[0109] Step 40264: Based on the feasible region of limb configuration, the allowable range of changes in the coordinates of the skeleton nodes at the next moment is defined, and the allowable range of changes is integrated with the envelope of the effective motion path to obtain the expected behavior constraint.
[0110] Optionally, the control ball, based on the feasible region of limb configuration preservation, uses this feasible region as a constraint to limit the allowable range of changes in the coordinates of all skeleton nodes of the effective tracking target at the next moment. The allowable range of changes refers to the pixel range within which the coordinates of the skeleton nodes can change position at the next moment, and this pixel range is entirely within the feasible region of limb configuration preservation. This allowable range of changes is then spatially integrated with the envelope of the effective motion path. Spatial range integration refers to the process of obtaining the joint spatial range of the allowable range of changes and the envelope of the effective motion path. The joint spatial range obtained after integration, which combines limb configuration constraints and motion path constraints, along with the relevant motion rules, is the expected behavior constraint.
[0111] Continuing with the above embodiment, the control ball retrieves the feasible region for maintaining the limb configuration of the right forearm. It then limits the allowable range of coordinate changes for the right elbow and right wrist endpoints of the effective tracking target at the next moment to a range of 300 to 330 pixels (horizontal coordinate) and 160 to 200 pixels (vertical coordinate) within this feasible region. Simultaneously, it retrieves the effective motion path envelope, which is a range of 60 to 390 pixels (horizontal coordinate) and 90 to 340 pixels (vertical coordinate). The control ball integrates this allowable coordinate change range with the effective motion path envelope to obtain a joint range that combines limb configuration and motion path constraints. This range represents the expected behavioral constraints for the right forearm.
[0112] The embodiments of the present invention further refine the logical constraint deduction process, realizing accurate prediction and multi-dimensional constraints on the movement behavior of the target being effectively tracked. This takes into account both the feasibility and safety of the target movement, as well as the rationality of the limb configuration, effectively avoiding tracking deviation and tracking loss caused by insufficient prediction of movement trends and failure to consider human body configuration constraints.
[0113] Optionally, the processes of steps 501 to 505 include: Step 501: Based on the displacement vector direction indicated by the instantaneous motion trend of the effectively tracked target and the activity boundary range defined by the expected behavior constraints, construct a dynamic search space with spatiotemporal topological constraints.
[0114] Optionally, the control ball extracts the instantaneous motion trend of the effectively tracked target. This instantaneous motion trend refers to the real-time motion features of the effectively tracked target in the current video frame, such as its motion direction, speed, and acceleration. It then extracts the corresponding displacement vector direction from this instantaneous motion trend. This displacement vector direction indicates the direction information of the effectively tracked target's possible positional movement in the next moment. Simultaneously, it retrieves expected behavior constraints and extracts the activity boundary range defined by these constraints. This activity boundary range refers to the reasonable spatial range set for the subsequent motion behavior and attitude changes of the effectively tracked target, serving as the constraint boundary for the target's motion. Subsequently, using the displacement vector direction as the core motion guide and the activity boundary range of the expected behavior constraints as the spatial limitation, a dynamic search space with both temporal and spatial topological constraints is constructed. The spatiotemporal topological constraint attribute means that the search space not only limits the target's spatial motion range but also, combined with the target's instantaneous motion trend, limits the temporal motion direction and amplitude. The dynamic search space refers to the image region range with spatiotemporal constraints delineated by the control ball to find the position of the effectively tracked target in the next frame.
[0115] In one embodiment, it is assumed that the instantaneous motion trend of the target being tracked by the control ball is a uniform upward motion to the right, with the corresponding displacement vector direction being 45 degrees upward to the right. The activity boundary range defined by the expected behavior constraint is a pixel area with horizontal coordinates from 200 to 400 and vertical coordinates from 100 to 300. The control ball uses the upward motion to the right at a 45-degree angle as the motion guide and uses the horizontal coordinates from 200 to 400 and the vertical coordinates from 100 to 300 as the spatial boundary to construct a dynamic search space with spatiotemporal topological constraints extending upward to the right at a 45-degree angle within this range.
[0116] Step 502: Based on the two-dimensional image coordinates of candidate pixels in the dynamic search space and the known three-dimensional spatial topology of the power grid high-altitude operation scenario, perform a reverse geometric mapping from the image plane to the real physical space to obtain a candidate physical location group.
[0117] Optionally, the monitoring sphere extracts the two-dimensional image coordinates of all candidate pixels within a dynamic search space, while simultaneously retrieving the known three-dimensional spatial topology of the power grid high-altitude operation scenario. Candidate pixels refer to all pixels within the dynamic search space that may contain the position of the target in the next frame; two-dimensional image coordinates refer to the pixel position coordinates with the upper left corner of the video frame image as the origin, the horizontal axis as the x-axis, and the vertical axis as the y-axis; the three-dimensional spatial topology refers to the three-dimensional geometric features of the actual physical spatial location, size, and structural shape of facilities such as poles and ladders in the power grid high-altitude operation scenario, serving as the benchmark for mapping the image plane to the real physical space. Then, based on preset camera imaging projection transformation rules, the two-dimensional image coordinates of the candidate pixels in the dynamic search space are subjected to a reverse geometric mapping from the image plane to the real physical space. This reverse geometric mapping refers to the spatial transformation process of restoring the two-dimensional image coordinates to three-dimensional coordinates in the real physical space, while maintaining the spatial topological relationship unchanged during the mapping process. All the mapped real physical space three-dimensional coordinates are integrated to form a candidate physical position group.
[0118] Continuing with the above embodiment, the control ball extracts the two-dimensional image coordinates {(250, 150), (260, 160), (270, 170)} of candidate pixels in the dynamic search space, and retrieves the three-dimensional spatial topology of the substation high-altitude operation scenario. This structure includes information such as the actual physical dimensions and spatial location of the tower. According to the camera imaging projection transformation rules, the control ball maps the two-dimensional image coordinates (250, 150) to the real physical space three-dimensional coordinates (5m, 3m, 8m), (260, 160) to (5.2m, 3.1m, 8.2m), and (270, 170) to (5.4m, 3.2m, 8.4m). These three-dimensional coordinates are then integrated to form candidate physical location groups {(5m, 3m, 8m), (5.2m, 3.1m, 8.2m), (5.4m, 3.2m, 8.4m)}.
[0119] Step 503: Based on the spatial distribution of each point in the candidate physical location group and the limb extension state represented by the local deformation descriptor of the effective tracking target, perform rigid body kinematic consistency verification to obtain physical trajectory segments that conform to human posture logic.
[0120] Optionally, the control ball analyzes the spatial distribution of all three-dimensional coordinate points within a candidate physical position group, while simultaneously retrieving the local deformation descriptor of the effective tracking target and extracting the limb extension state represented by the descriptor. The spatial distribution refers to the spatial characteristics of the candidate physical position points in real physical space, such as their arrangement, spacing, and extension direction; the limb extension state refers to the limb geometric characteristics of the effective tracking target, such as the current degree of limb opening, the relative positions of each segment, and the arrangement of skeletal nodes. Based on the fundamental laws of rigid body kinematics, a rigid body kinematic consistency check is performed on the spatial distribution of the candidate physical position group and the limb extension state of the effective tracking target. This checks whether the spatial changes of the candidate physical positions conform to the motion laws of the human limb as a rigid body and whether they match the current limb extension state. Specific checks include whether the amplitude and direction of limb movement are consistent with the displacement and rotation laws of rigid body movement and whether they conform to the motion limits of the human limb in its extension state. Candidate physical locations that fail the verification are removed, and the candidate physical locations that pass the verification are sorted according to the direction of movement to form a physical trajectory segment that conforms to the logic of human posture. That is, the physical trajectory segment refers to the short-distance movement trajectory of the effective tracking target in real physical space that conforms to the law of rigid body movement of the human body.
[0121] Continuing with the above embodiment, the spatial distribution of the candidate physical location groups {(5m, 3m, 8m), (5.2m, 3.1m, 8.2m), and (5.4m, 3.2m, 8.4m)} analyzed by the control ball is a uniform extension along the upper right, with consistent spacing. The limb extension state represented by the local deformation descriptor is a climbing posture with arms extended and legs slightly bent. The control ball performs a consistency check based on the laws of rigid body kinematics, determining that the spatial distribution pattern matches the rigid body motion laws of the human limbs in a climbing posture, with no cases exceeding the limits of limb movement; the check passes. The control ball sorts these candidate physical locations according to the upper right movement direction, forming physical trajectory segments that conform to the logic of human posture.
[0122] Step 504: Based on the positional continuity and temporal change characteristics of the physical trajectory fragments in the continuous time series, perform cross-frame spatial connectivity inverse geometric mapping to obtain the reconstructed three-dimensional motion path.
[0123] Optionally, the control ball analyzes the positional continuity and temporal variation characteristics of physical trajectory segments over a continuous time series. Positional continuity refers to the absence of abrupt changes or breaks in the spatial position of each point within the physical trajectory segment over a continuous time period. Temporal variation characteristics refer to the changing patterns of the velocity and direction of motion of each point within the physical trajectory segment over a continuous time period. Subsequently, a cross-frame spatial connectivity reverse geometric mapping is performed on the physical trajectory segments. This cross-frame spatial connectivity reverse geometric mapping refers to the spatial transformation process of combining a single physical trajectory segment with image information from multiple consecutive frames to construct a continuous and connected long-distance motion trajectory in real physical space. Specifically, the process involves: arranging the segments in chronological order according to the continuous time series; calculating the positional distance between adjacent physical trajectory segments in the time series; determining whether the distance falls within a preset continuity threshold; if within the threshold, smoothly connecting the two trajectory segments through reverse geometric mapping and supplementing the physical position points at the connection point; if outside the threshold, determining the trajectory segment as invalid and discarding it; repeating the above process to connect and integrate all continuous physical trajectory segments to form a complete, continuous, and uninterrupted three-dimensional motion path. In other words, the three-dimensional motion path refers to the complete motion trajectory of the target in real physical space that conforms to the laws of human movement and is continuously connected across frames.
[0124] Continuing with the above embodiment, the control ball extracts physical trajectory segments from a single frame {(5m, 3m, 8m), (5.2m, 3.1m, 8.2m), (5.4m, 3.2m, 8.4m)}, analyzing their positional continuity as uninterrupted and with uniform spacing, and their temporal variation as uniform upward and rightward movement. The control ball performs cross-frame spatial connectivity reverse geometric mapping on these physical trajectory segments across 5 consecutive frames, connecting the beginning and end of the physical trajectory segments in each frame to form a continuous trajectory extending from (5m, 3m, 8m) to (6m, 3.5m, 9m), thus obtaining the reconstructed three-dimensional motion path.
[0125] Step 505: Based on the safety operation procedure logic in the three-dimensional motion path and expected behavior constraints, perform spatial compliance consistency verification to obtain the target tracking result.
[0126] Optionally, the monitoring sphere, based on the three-dimensional motion path, retrieves the safety operation procedure logic contained in the expected behavior constraints. This safety operation procedure logic refers to the safety rules that the target motion path must comply with, as set according to the power grid high-altitude operation safety regulations. These rules include that the motion path must not exceed the safe operation area and must not contact dangerous facilities. The three-dimensional motion path is then checked for spatial compliance consistency with the safety operation procedure logic. This involves determining whether the three-dimensional motion path conforms to the power grid high-altitude operation safety regulations. Specific checks include whether the motion path is within the safe operation area, whether there is any contact with dangerous facilities, and whether it complies with the high-altitude operation motion safety regulations. Finally, the verified three-dimensional motion path is integrated with the effective tracking target's position coordinates and attitude changes in continuous video frames to form a complete and accurate set of target motion information, which is the target tracking result. This target tracking result includes the complete three-dimensional motion path, motion status, and compliance indicators of the effectively tracked target.
[0127] Continuing with the above embodiment, the 3D motion path extracted and reconstructed by the monitoring ball is a continuous trajectory from (5m, 3m, 8m) to (6m, 3.5m, 9m). The ball retrieves the safety operation procedure logic from the expected behavior constraints. This logic stipulates that the operator's movement path must remain within a 1-meter safe working area around the tower ladder and must not touch any live parts of the tower. The monitoring ball performs a spatial compliance consistency check, determining that the 3D motion path is entirely within the 1-meter safe working area around the ladder and has not touched any live parts; the check passes. The monitoring ball integrates this 3D motion path with the effective tracking target's position coordinates, climbing posture changes, and other information over five consecutive frames to form a complete target motion information set, which is the target tracking result.
[0128] The embodiments of the present invention improve the entire process of target tracking, effectively solve the problems of tracking offset and tracking loss that occur in existing methods due to their reliance on fixed global feature templates in scenarios such as complex backgrounds, small targets, and changes in target features. It improves the accuracy and stability of target tracking in high-altitude power grid operations, fully presents the real motion state of the effectively tracked target, and provides a reliable basis for safety monitoring.
[0129] Furthermore, the target tracking device for high-altitude operations in power grids provided by the present invention will be described below. The target tracking device for high-altitude operations in power grids described below can be referred to in correspondence with the target tracking method for high-altitude operations in power grids described above.
[0130] Optional, refer to Figure 2 , Figure 2 This is a schematic diagram of the power grid high-altitude operation target tracking device provided by the present invention. The power grid high-altitude operation target tracking device includes: The standardization reconstruction module 210 performs temporal difference analysis on continuous frame images from the power grid high-altitude operation video stream to obtain candidate target image patches containing potential targets. It then performs topological standardization reconstruction based on the foreground pixel distribution density of the candidate target image patches to obtain standard-scale image patches under a unified geometric benchmark. The target image enhancement module 220 performs high-frequency detail enhancement and edge contour sharpening reconstruction based on the texture gradient features of the standard-scale image patches to obtain enhanced target image patches. The deformation descriptor construction module 230 constructs a local deformation descriptor representing the geometric features of the target posture based on the skeleton node coordinates extracted from the enhanced target image patches. The behavior constraint deduction module 240 verifies the rationality of the human body structure based on the relative topological relationship of the nodes in the local deformation descriptor to obtain an effective tracking target. It then performs logical constraint deduction based on the spatial relationship between the effective tracking target and the surrounding environment, combined with the temporal changes of the local deformation descriptor, to obtain expected behavior constraints. The dynamic search verification module 250 constructs a dynamic search space based on the instantaneous motion trend of the effective tracking target and expected behavior constraints, and performs inverse geometric mapping and consistency verification in the dynamic search space to obtain the target tracking result.
[0131] This invention addresses the issues of tracking offset and tracking loss that arise in existing methods due to their reliance on fixed global feature templates in scenarios with complex backgrounds, small targets, and changing target features. It improves the accuracy and stability of target tracking in high-altitude power grid operations, thus meeting the requirements for high-precision safety monitoring of high-altitude power grid operations.
[0132] Please see Figure 3 , Figure 3 This is an embodiment diagram of a control ball provided in this invention. Figure 3 As shown, this embodiment of the invention provides a control ball 300, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, it implements the following steps 10 to 50.
[0133] Please see Figure 4 , Figure 4 An embodiment diagram of a computer-readable storage medium provided in accordance with an embodiment of the present invention is shown. Figure 4 As shown, this embodiment provides a computer-readable storage medium 400 on which a computer program 311 is stored. When the computer program 311 is executed by a processor, it performs the following steps 10 to 50.
[0134] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the power grid high-altitude operation target tracking method provided by the above methods, which includes the process of steps 10 to 50.
[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for tracking targets during high-altitude operations in power grids, characterized in that, include: Temporal difference analysis was performed on continuous frame images of high-altitude power grid operation video stream to obtain candidate target image patches containing potential targets. Based on the foreground pixel distribution density of the candidate target image patches, topological normalization reconstruction was performed to obtain standard-scale image patches under a unified geometric benchmark. High-frequency detail enhancement and edge contour sharpening reconstruction are performed based on the texture gradient features of standard-scale image patches to obtain target enhanced image patches; Based on the skeleton node coordinates extracted from the target augmented image patch, a local deformation descriptor representing the geometric features of the target pose is constructed; The rationality of human body structure is verified based on the relative topological relationship of nodes in the local deformation descriptor, and an effective tracking target is obtained. Logical constraint deduction is then performed based on the spatial relationship between the effective tracking target and the surrounding environment, combined with the temporal changes of the local deformation descriptor, to obtain the expected behavior constraint. A dynamic search space is constructed based on the instantaneous motion trend and expected behavior constraints of the target for effective tracking. Inverse geometric mapping and consistency verification are performed in the dynamic search space to obtain the target tracking result.
2. The target tracking method for high-altitude operations in power grids according to claim 1, characterized in that, The topological normalization reconstruction based on the foreground pixel distribution density of candidate target image patches yields standard-scale image patches under a unified geometric benchmark, including: Calculate the geometric centroid coordinates of all foreground pixels based on the two-dimensional coordinates of all foreground pixels in the candidate target image block, and determine the geometric centroid coordinates as the topological center reference point of the candidate target image block; Based on the topological center reference point, rays are emitted in the four boundary directions of the candidate target image block and the farthest intersection point of the rays with the foreground pixel point set is detected to obtain a four-way extreme boundary point set including the farthest intersection point of the upper boundary, the farthest intersection point of the lower boundary, the farthest intersection point of the left boundary, and the farthest intersection point of the right boundary. Based on the vertical Euclidean distance between the farthest intersection point of the upper boundary and the farthest intersection point of the lower boundary in the set of four-way extreme boundary points, the vertical topological span of the candidate target image block is determined, and based on the horizontal Euclidean distance between the farthest intersection point of the left boundary and the farthest intersection point of the right boundary in the set of four-way extreme boundary points, the horizontal topological span of the candidate target image block is determined. The first geometric scaling ratio in the vertical direction and the second geometric scaling ratio in the horizontal direction are calculated based on the preset fixed height and preset fixed width combined with the vertical topological span and the horizontal topological span. Based on the topological center reference point of the candidate target image patch, topological mapping is performed by combining the first geometric scaling ratio and the second geometric scaling ratio to obtain the standard scale image patch.
3. The target tracking method for high-altitude operations in power grids according to claim 2, characterized in that, The standard-scale image patch is obtained by performing topological mapping based on the topological center reference point of the candidate target image patch and combining the first geometric scaling ratio and the second geometric scaling ratio, including: Based on the topology center reference point of the candidate target image patch, a standard topology mapping grid covering the entire area of the candidate target image patch is constructed using the first geometric scaling ratio and the second geometric scaling ratio. Based on the position of each foreground pixel in the candidate target image block in the standard topological mapping grid, a direct coordinate mapping relationship table for the foreground pixels of the candidate target image block is established. Based on the direct coordinate mapping table, the grayscale or color values of the foreground pixels in each grid cell of the candidate target image block are filled into the corresponding target pixel positions in the standard scale image to obtain the initial scale image block for foreground texture information topological transfer. Based on the initial scale image block, the remaining blank pixel positions without foreground texture information are filled with default background values to obtain the standard scale image block.
4. The target tracking method for high-altitude operations in power grids according to claim 1, characterized in that, The method of verifying the rationality of human body structure based on the relative topological relationship of nodes using local deformation descriptors to obtain effective tracking targets includes: Based on the coordinates of all skeleton nodes and their node type identifiers in the local deformation descriptor, directly connected skeleton node pairs are extracted to obtain the initial limb connection group. Based on the coordinates of the two endpoint skeleton nodes of each limb connection in the initial limb connection group, the set of length ratios between the length of each limb connection and the reference length is calculated, and the set of length ratios is compared one by one with the preset reasonable proportion range of human anatomy to screen out qualified limb connection groups that pass the length ratio verification. Based on the coordinates of three consecutive connected skeleton nodes in the qualified limb connection group, a joint angle geometry structure composed of two adjacent limbs is constructed. The internal angle value of each joint angle geometry structure is compared with the preset human joint physiological activity angle threshold to screen out qualified joint structure groups that pass the angle verification. Based on all skeleton nodes in the qualified joint structure group, the connection paths between nodes are reconstructed, and isolated skeleton nodes that are disconnected or floating limb branches that are not connected to the main node are removed to obtain a connected skeleton node group with complete topological connectivity. Based on the connected skeleton node group, the original local deformation descriptor is retrieved in reverse, mismatched node data and connection relationships are located and removed to obtain the remaining node data and connection relationships, and the remaining node data and connection relationships are re-encapsulated to obtain the effective tracking target.
5. The target tracking method for high-altitude power grid operations according to claim 4, characterized in that, The logical constraint deduction based on the spatial relationship between the effective tracking target and the surrounding environment, combined with the temporal changes of local deformation descriptors, yields the expected behavioral constraints, including: Based on the skeleton node coordinate sequence in the local deformation descriptor, the relative displacement vector direction of the skeleton node coordinates in consecutive frames is extracted to obtain the limb movement tendency vector group. Based on the vector direction with the largest magnitude in the limb movement tendency vector group, the main axis direction of limb extension is identified to obtain the locked potential action force point. Based on the geometric center position calculated by combining the trunk center node and bilateral hip joint nodes of the effective tracking target with the potential action force point, an equivalent balance reference point is obtained. Then, based on the coordinates of the pole structure surface in the surrounding environmental topology elements and the coordinates of the end skeleton node of the effective tracking target, a spatial proximity projection is performed to obtain a candidate environmental contact point group. Based on the candidate environmental contact point group, a dynamic support topology is constructed with the minimum convex hull boundary. It is then determined whether the equivalent equilibrium reference point falls inside the dynamic support topology domain to obtain the target's attitude stability state. Based on the attitude stability state and the normal direction of the tower structure surface in the surrounding environmental topology, the projection components of the equivalent balance reference point and the potential action force point on the normal direction of the tower structure surface are calculated to obtain the attitude instability trend vector. Based on the extension direction of the attitude instability trend vector, the regions where the surface normal direction and the attitude instability trend vector have an inverse blocking topological relationship are screened to obtain environmental contact candidate regions. Spatial proximity matching is performed between the candidate environmental contact regions and the coordinates of the end skeleton nodes in the local deformation descriptor to obtain the actual contact feature node pairs between the target and the environment, and the physical anchoring point of the target in the environmental topology is determined based on the actual contact feature node pairs. Based on the physical anchor points and the limb movement tendency vector group, geometric constraint deduction is performed to obtain the expected behavior constraint.
6. The target tracking method for high-altitude operations in power grids according to claim 5, characterized in that, The geometric constraint deduction based on the physical anchor points and the limb movement tendency vector group yields the expected behavioral constraints, including: Based on the physical anchor points and the limb motion tendency vector group, the motion trajectory of the remaining degrees of freedom after being constrained by the physical anchor points is derived to obtain the restricted motion path envelope. Based on the restricted motion path envelope, invalid motion components that exceed the physical boundary of the tower structure are eliminated to obtain the effective motion path envelope. Based on the effective motion path envelope and the preset safe operation area topology range, a spatial inclusion comparison is performed to obtain a path compliance identifier. Based on the path compliance identifier, the critical motion segment that is about to cross the safety boundary is marked to obtain the critical marked motion segment. Based on the critical marker motion segment and the physical anchor point, a rigid motion trajectory ring is constructed with the physical anchor point as the rotation center and the preset anatomical length of the corresponding limb segment as the radius. Based on the preset physiological activity sector of the human joint, the rigid motion trajectory ring and the preset physiological activity sector are geometrically intersected to obtain the limb configuration preservation feasible region. Based on the limb configuration, the feasible domain is defined to limit the allowable range of changes in the skeletal node coordinates at the next moment. The allowable range of changes is then integrated with the effective motion path envelope to obtain the expected behavioral constraint.
7. The target tracking method for high-altitude operations in power grids according to claim 1, characterized in that, The steps for obtaining target tracking results include: Based on the displacement vector direction indicated by the instantaneous motion trend of the effective tracking target and the activity boundary range defined by the expected behavior constraints, a dynamic search space with spatiotemporal topological constraints is constructed. Based on the two-dimensional image coordinates of candidate pixels in the dynamic search space and the known three-dimensional spatial topology of the power grid high-altitude operation scenario, a reverse geometric mapping from the image plane to the real physical space is performed to obtain a candidate physical location group. Based on the spatial distribution of each point in the candidate physical location group and the limb extension state represented by the local deformation descriptor of the effective tracking target, a rigid body kinematic consistency check is performed to obtain a physical trajectory segment that conforms to the logic of human posture. Based on the positional continuity and temporal variation characteristics of the physical trajectory segments in a continuous time series, a cross-frame spatial connectivity inverse geometric mapping is performed to obtain the reconstructed three-dimensional motion path; Based on the spatial compliance consistency verification between the three-dimensional motion path and the safety operation procedure logic in the expected behavior constraints, the target tracking result is obtained.
8. A target tracking device for high-altitude operations in power grids, characterized in that, Applied to the power grid high-altitude operation target tracking method as described in any one of claims 1 to 7; The power grid high-altitude operation target tracking device includes: The standardization reconstruction module is used to perform temporal difference analysis on continuous frame images of high-altitude power grid operation video stream to obtain candidate target image blocks containing potential targets, and to perform topological standardization reconstruction based on the foreground pixel distribution density of the candidate target image blocks to obtain standard-scale image blocks under a unified geometric benchmark. The target image enhancement module performs high-frequency detail enhancement and edge contour sharpening reconstruction based on the texture gradient features of standard-scale image patches to obtain target enhanced image patches; The deformation descriptor construction module is used to construct a local deformation descriptor that represents the geometric features of the target pose based on the skeleton node coordinates extracted from the target augmented image patch. The behavior constraint deduction module is used to verify the rationality of human body structure based on the relative topological relationship of nodes in the local deformation descriptor, obtain the effective tracking target, and perform logical constraint deduction based on the spatial relationship between the effective tracking target and the surrounding environment and the temporal change of the local deformation descriptor to obtain the expected behavior constraint. The dynamic search and verification module is used to construct a dynamic search space based on the instantaneous motion trend and expected behavior constraints of the effectively tracked target, and to perform reverse geometric mapping and consistency verification in the dynamic search space to obtain the target tracking result.
9. A control ball, comprising: Memory, used to store computer software programs; A processor for reading and executing the computer software program, characterized in that, when the processor executes the computer software program, it implements the power grid high-altitude operation target tracking method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that, The storage medium stores a computer software program, which, when executed by a processor, implements the power grid high-altitude operation target tracking method as described in any one of claims 1 to 7.