A power grid operation scene personnel target real-time tracking method, system, device and medium

The PTRTT-Net network solves the accuracy and robustness issues of target tracking in power grid operation scenarios through multi-scale enhancement and background suppression mechanisms, combined with group feature interaction and lightweight network design, and achieves efficient real-time tracking effects.

CN120355942BActive Publication Date: 2025-10-10GUIZHOU POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510830263.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-10
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Traditional target tracking methods are difficult to adapt to complex environments in power grid operation scenarios, resulting in target loss, mistracking, false detection and missed detection, especially in high-dynamic and complex scenarios where real-time and robustness are insufficient.

Method used

The PTRTT-Net network is adopted to achieve real-time tracking of human targets in power grid operation scenarios through multi-scale enhancement and background suppression mechanism, combined with group feature interaction and lightweight network design.

Benefits of technology

The accuracy and robustness of personnel target tracking in power grid operation scenarios are improved, computational overhead is reduced, and high-precision real-time tracking performance is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355942B_ABST
    Figure CN120355942B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image recognition, and discloses a power grid operation scene personnel target real-time tracking method, system, device and medium, the method comprising: acquiring a first data set under a target power grid operation scene, and performing first preprocessing on the first data set to obtain a second data set; establishing a first tracking model, taking the second data set as a training set and a verification set of the first tracking model; and performing target power grid operation scene personnel target real-time tracking according to the first tracking model after verification is completed. The present application solves the target false detection and missed detection problems caused by background noise and illumination changes in the prior art; the method effectively highlights the key features of the personnel target through a multi-scale foreground enhancement and background suppression mechanism, while suppressing background interference, so that personnel target tracking in a complex power grid operation scene is more stable and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a method, system, device and medium for real-time tracking of personnel targets in power grid operation scenarios. Background Art

[0002] Power grid operations, as a complex industrial environment, are often characterized by large-scale equipment, dense personnel activity, and diverse operational scenarios. Real-time tracking of personnel in these scenarios is not only a crucial technical approach to ensuring operational safety and improving efficiency, but also a key component of smart grid management and the intelligent upgrade of industrial operations. However, traditional methods for tracking personnel targets in power grid operations are significantly limited by their openness, dynamism, and complexity. New technologies are urgently needed to improve the accuracy and robustness of real-time tracking.

[0003] Target tracking in power grid operation scenarios faces multiple technical challenges. On the one hand, the operation site environment is often significantly complex, with features such as multiple target occlusions, dramatic lighting changes, and strong background interference. Traditional feature engineering-based methods (such as color segmentation and shape matching) have difficulty effectively distinguishing between target and background information. On the other hand, due to the high level of activity of human targets in power grid operation environments, often accompanied by large posture changes and frequent irregular movements, traditional target tracking methods are difficult to adapt to the needs of complex scenarios and are prone to target loss and mistracking. In addition, power grid operation sites have high real-time requirements, requiring rapid and stable target tracking in a dynamically changing environment, which places higher demands on the efficiency and robustness of the algorithm.

[0004] Traditional target tracking methods, such as those based on optical flow, detection-based tracking (DBT), and Kalman filtering, while effective in some simple scenarios, struggle to meet the demands of highly dynamic and complex scenarios like power grid operations. These methods often rely on handcrafted features and lack adaptability to target appearance, scale variations, and environmental complexity. This makes it difficult to ensure real-time performance and stability, especially in the presence of occlusion and rapid target movement. Summary of the Invention

[0005] In view of the above existing problems, the present invention is proposed.

[0006] Therefore, the present invention provides a method, system, device and medium for real-time tracking of personnel targets in power grid operation scenarios, which can solve the accuracy and robustness problems of target tracking in power grid operation scenarios, as well as the problems of target misdetection and missed detection caused by background noise and lighting changes.

[0007] To solve the above technical problems, the present application provides the following technical solutions:

[0008] In a first aspect, the present application provides a power grid operation scene personnel target real-time tracking method, comprising:

[0009] Obtaining a first data set under a target power grid operation scene, and performing first preprocessing on the first data set to obtain a second data set;

[0010] Establishing a first tracking model, taking the second data set as the training set and the verification set of the first tracking model;

[0011] The first tracking model comprises a first enhancement and suppression module, a second grouped feature interaction module, and a third target tracking module;

[0012] The first enhancement and suppression module comprises a multi-scale enhancement attention submodule, and the first enhancement and suppression module inputs the second feature map and the seventh feature map into the second grouped feature interaction module;

[0013] The multi-scale enhancement attention submodule is used to generate a second feature map according to the second data set, and the output of the first enhancement and suppression module is a seventh feature map;

[0014] The second grouped feature interaction module is used to output image detection frame information;

[0015] According to the first tracking model after verification, the target power grid operation scene personnel target real-time tracking is performed.

[0016] As a preferred scheme of the power grid operation scene personnel target real-time tracking method of the present application, wherein: the second feature map generated according to the second data set comprises:

[0017] Three paths are established, and three different feature maps are obtained according to the second data set;

[0018] The three paths correspond to maximum pooling operations of different sizes;

[0019] The three different feature maps are subjected to a padding operation, and the sizes of the three different feature maps after the padding operation are the same.

[0020] As a preferred scheme of the power grid operation scene personnel target real-time tracking method of the present application, wherein: the second feature map generated according to the second data set further comprises:

[0021] The second data set is subjected to several convolution operations and up-sampling operations to obtain another feature map with the same size as the three different feature maps;

[0022] Performing a splicing operation on the three different feature maps and another feature map;

[0023] A first feature enhancement extraction operation is performed on the feature map after the splicing operation to obtain a second feature map.

[0024] As a preferred solution of the method for real-time tracking of personnel targets in power grid operation scenarios according to the present invention, the first feature enhancement and extraction operation includes:

[0025] Performing two convolution operations of different sizes on the feature map after the splicing operation, and recording the results as a first convolution feature map and a second convolution feature map;

[0026] Generating a target enhancement weight for the second convolutional feature map through an activation function;

[0027] The target enhancement weight is multiplied by the first convolution feature map element by element, and the element-wise multiplication result is upsampled to obtain a second feature map.

[0028] As a preferred solution of the method for real-time tracking of personnel targets in power grid operation scenarios according to the present invention, the first preprocessing of the first data set includes:

[0029] Annotate the first dataset, use a video annotation tool to select the target position of the person in the video, and annotate the label;

[0030] The frame selection operation is to determine the horizontal and vertical coordinates of the upper left corner and the lower right corner of the frame in each frame;

[0031] The label annotation includes the target person's label name, annotation box information, and the target's trajectory changing over time in the video;

[0032] The first data set is a video data set.

[0033] As a preferred solution of the method for real-time tracking of personnel targets in power grid operation scenarios according to the present invention, wherein: the second grouping feature interaction module includes a first splitting operation;

[0034] The first splitting operation is used to split the feature map into a plurality of sub-maps;

[0035] Each subgraph is equipped with two subgraph processing paths, one of which is used to generate intra-group weights, and the other is used to generate interactive fusion feature maps.

[0036] As a preferred scheme of the power grid operation scene personnel target real-time tracking method, the third target tracking module comprises three different branches, which are divided into a first branch, a second branch and a third branch, the first branch is used for extracting compact features in the image detection box information, the second branch is used for extracting local deep features in the image detection box information, and the third branch is used for extracting lightweight features in the image detection box information.

[0037] In a second aspect, the present application provides a power grid operation scene personnel target real-time tracking system, comprising:

[0038] A data acquisition and processing module is configured to acquire a first data set in a target power grid operation scene, and perform first preprocessing on the first data set to obtain a second data set.

[0039] A model establishing module is configured to establish a first tracking model, and use the second data set as a training set and a verification set of the first tracking model.

[0040] The first tracking model comprises a first enhancement and suppression module, a second grouped feature interaction module and a third target tracking module.

[0041] The first enhancement and suppression module comprises a multi-scale enhanced attention submodule, and the first enhancement and suppression module inputs a second feature map and a seventh feature map into the second grouped feature interaction module.

[0042] The multi-scale enhanced attention submodule is configured to generate a second feature map according to the second data set, and the output of the first enhancement and suppression module is a seventh feature map.

[0043] The second grouped feature interaction module is configured to output image detection box information.

[0044] A tracking module is configured to perform target real-time tracking of personnel in a target power grid operation scene according to the first tracking model after verification.

[0045] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method when executing the computer program.

[0046] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the method when executed by a processor.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention proposes a method, system, device and medium for real-time tracking of personnel targets in power grid operation scenarios, obtains a first data set in the target power grid operation scenario, and performs a first preprocessing on the first data set to obtain a second data set; establishes a first tracking model, uses the second data set as the training set and verification set of the first tracking model; and performs real-time tracking of personnel targets in the target power grid operation scenario based on the first tracking model after verification. The present invention solves the problems of false detection and missed detection of targets caused by background noise and illumination changes in the prior art; the method effectively highlights the key features of personnel targets through a multi-scale foreground enhancement and background suppression mechanism, while suppressing background interference, making personnel target tracking in complex power grid operation scenarios more stable and reliable; in addition, the method significantly improves the feature extraction efficiency and real-time performance in multi-target dense scenarios through grouped feature interaction optimization and lightweight network design, reduces the computational overhead and insufficient real-time performance caused by complex network structures in the prior art, thereby achieving high-precision real-time tracking performance while ensuring lightweight. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0049] Figure 1 A method flow chart of a method for real-time tracking of personnel targets in power grid operation scenarios provided by one embodiment of the present invention.

[0050] Figure 2 A schematic diagram of the overall structure of PTRTT-Net, a method for real-time tracking of personnel targets in power grid operation scenarios provided by one embodiment of the present invention.

[0051] Figure 3 This is a structural diagram of a multi-scale foreground enhancement and background suppression module for a real-time tracking method for personnel targets in power grid operation scenarios provided by one embodiment of the present invention.

[0052] Figure 4 This is a structural diagram of the MSEA submodule of a method for real-time tracking of personnel targets in power grid operation scenarios provided by one embodiment of the present invention.

[0053] Figure 5 This is a structural diagram of a grouping feature interaction module of a method for real-time tracking of personnel targets in power grid operation scenarios provided by one embodiment of the present invention.

[0054] Figure 6A lightweight extensible target tracking module structure diagram of a power grid operation scene personnel target real-time tracking method is provided for an embodiment of the present application.

[0055] Figure 7 An internal structure diagram of a computer device of a power grid operation scene personnel target real-time tracking method is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.

[0057] Embodiment 1, with reference to Figure 1-Figure 7 For the first embodiment of the present application, the embodiment provides a power grid operation scene personnel target real-time tracking method, comprising:

[0058] Before detailing the embodiments of the present application, for the sake of clarity, some related concepts are first explained.

[0059] VIA (VGG Image Annotator) is a lightweight open-source image and video annotation tool. It runs based on a browser, does not need to be installed, and can be used offline. VIA supports multiple annotation types such as rectangles, points, lines, polygons, and text labels, and is suitable for tasks such as object detection and image segmentation. Its simple and intuitive interface and highly customizable functions make it a commonly used data annotation tool in machine learning and computer vision projects.

[0060] Image frame: An image frame is the smallest unit that makes up a video, and is essentially an image.

[0061] CBS operation is a basic convolutional module that extracts local features of an image through a 3x3 convolution kernel, then performs BatchNorm to stabilize the feature distribution, and introduces nonlinearity through an activation function to enhance the feature expression ability of the model.

[0062] The Detection detection head is a key component of the target detection model, which is used to predict the class and bounding box of the target from the feature map generated by the feature extraction network.

[0063] The Kalman filter is a recursive algorithm used to estimate the target state from continuous observations of a dynamic system. It is widely used in target tracking and navigation, providing smooth and accurate state predictions in the presence of noise. The Hungarian algorithm is an optimization method used to efficiently solve the assignment problem, achieving the best match by minimizing the total cost in a cost matrix. The two are often combined in multi-target tracking, with the Kalman filter used for state prediction and the Hungarian algorithm used for matching targets to observations.

[0064] PTRTT-Net (personnel targets Real time tracking network): A real-time tracking network for personnel targets in power grid operation scenarios, i.e., the network used by the first tracking model in this application, uses PTRTT-Net to perform real-time tracking of personnel targets in power grid operation scenarios.

[0065] MSEA (Multi-Scale Enhance Attention): Multi-scale enhanced attention module, that is, the multi-scale enhanced attention submodule in this application,

[0066] There are some problems in the existing related technologies, such as the large dynamic changes of personnel targets in power grid operation scenarios, which makes the targets easily lost or mistracked during the tracking process, as well as the misdetection and missed detection of targets due to background noise and lighting changes.

[0067] This application provides a method that can effectively solve the above-mentioned problems. Next, we will combine multiple embodiments to elaborate on how to implement the real-time tracking method of personnel targets in power grid operation scenarios;

[0068] Figure 1 A method for real-time tracking of personnel targets in power grid operation scenarios is shown, including:

[0069] S101, obtaining a first data set in a target power grid operation scenario, and performing a first preprocessing on the first data set to obtain a second data set;

[0070] In an optional embodiment, the target power grid operation scenarios may include substations, transmission line inspections, distribution room maintenance, and other scenarios. These scenarios often feature complex and changing backgrounds, varying lighting conditions, and potential human occlusion. To effectively track human operators in these scenarios, relevant video data must first be collected as a first dataset. These datasets should cover a wide range of possible operational scenarios to ensure the broad applicability of the trained tracking model.

[0071] It should be noted that the first preprocessing step is to improve data quality and model training efficiency. Preprocessing steps may include, but are not limited to, data cleaning, format conversion, and normalization to ensure that the first dataset meets the requirements of model training.

[0072] In an embodiment of the present application, performing first preprocessing on the first data set includes:

[0073] Label the first dataset, use the video annotation tool to select the target position of the person in the video, and label it;

[0074] The frame selection operation is to determine the horizontal and vertical coordinates of the upper left corner and lower right corner of the frame in each frame;

[0075] The label annotation includes the target person's label name, annotation box information, and the target's trajectory changing over time in the video;

[0076] The first dataset is a video dataset.

[0077] Specifically, in this application, the specific steps of performing the first preprocessing on the first data set are as follows:

[0078] Construct an original video dataset for real-time tracking of personnel targets in power grid operation scenarios, or use an existing publicly available video dataset of power grid operation scenarios, i.e., the first dataset. Collect video data of personnel activities in the power grid operation environment (including but not limited to field shooting at different times and locations through surveillance cameras and drones, or using existing publicly available video datasets), and manually annotate the collected video data. Use a video annotation tool to select the position of the personnel target in the video (i.e., determine the horizontal and vertical coordinates of the upper left and lower right corners of the frame in each frame) and label them. The labels include the target person's label name, the annotation box information, and the target's trajectory over time in the video.

[0079] Furthermore, the original video of the real-time monitoring of personnel targets in the power grid operation scene is preprocessed to obtain a preprocessed video frame image dataset;

[0080] Furthermore, the data set is divided and subsequently used to train a real-time tracking model for personnel targets in power grid operation scenarios.

[0081] Exemplarily, the specific first pre-processing process for the original video is:

[0082] Image frames are extracted frame by frame from the original video data and uniformly adjusted to a fixed resolution to ensure consistent image frame sizes for subsequent model training and testing.

[0083] It should be noted that due to the limited amount of data in the original video, in order to ensure that the dataset has sufficient samples for training, verification and testing, and to enhance the robustness and adaptability of the target tracking model, the image frames are then data enhanced to generate a data-enhanced video frame image dataset.

[0084] Furthermore, specific data enhancement methods include random rotation, horizontal flipping, illumination change simulation, adding motion blur and color jitter, etc., to obtain the second data set.

[0085] Furthermore, the data-enhanced video frame image dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0086] It should be noted that obtaining a first data set under the target power grid operation scenario and performing a first preprocessing on the first data set to obtain a second data set can improve the quality and diversity of the data and provide a solid foundation for subsequent model training. Accurately marking the target positions of personnel through video annotation tools can ensure that the model learns key target features during the training process. At the same time, data enhancement processing can simulate lighting changes and personnel movements in various actual situations, further enhancing the generalization ability of the model. In addition, dividing the data set into training set, validation set and test set helps to conduct effective performance evaluation and parameter adjustment during the model training process, ensuring that the final tracking model has excellent real-time tracking performance and robustness.

[0087] S102, establishing a first tracking model, and using the second data set as a training set and a validation set of the first tracking model;

[0088] In an optional embodiment, the first tracking model can be built using a deep learning framework, such as TensorFlow or PyTorch. The model uses an end-to-end training approach and can directly predict the location information of the operator from the input power grid operation scene video frames.

[0089] In an optional embodiment, the first tracking model can also be built using machine learning, for example, classic machine learning algorithms such as support vector machines (SVMs) and random forests, to track targets using manually extracted features. However, compared to deep learning models, their tracking accuracy and generalization capabilities may be slightly inferior.

[0090] It should be noted that the first tracking models established by the above methods have certain limitations. For example, although the deep learning model has high tracking accuracy, the model complexity is high, the calculation amount is large, and the hardware resource requirements are high; and although the machine learning model has a small amount of calculation, the manual feature extraction method may not be able to fully capture the key information of the target, resulting in limited tracking accuracy.

[0091] To overcome the above limitations, the present application proposes a PTRTT-Net, as shown in Figure 2 The model adopts a multi-scale foreground enhancement and background suppression mechanism, effectively highlighting the key features of personnel targets while suppressing background interference. PTRTT-Net significantly improves the feature extraction efficiency and real-time performance in multi-target dense scenes through grouped feature interaction optimization and lightweight network design. Specifically, the architecture of PTRTT-Net embeds a multi-scale enhanced attention module (MSEA), which can adaptively enhance target features and suppress background noise, thereby improving the stability and accuracy of tracking. In addition, PTRTT-Net also adopts a grouped feature interaction module, which further optimizes feature representation and improves the tracking performance of the model through grouping processing and feature interaction. Lightweight network design ensures that PTRTT-Net maintains high-precision tracking while achieving low computational overhead and high real-time performance, making it suitable for real-time tracking tasks in complex power grid operation scenarios.

[0092] In an embodiment of the present application, the first tracking model includes a first enhancement and suppression module, a second grouped feature interaction module, and a third target tracking module.

[0093] In an embodiment of the present application, the first enhancement and suppression module includes a multi-scale enhanced attention submodule, and the first enhancement and suppression module inputs the second feature map and the seventh feature map into the second grouped feature interaction module.

[0094] In an embodiment of the present application, the multi-scale enhanced attention submodule is used to generate the second feature map according to the second data set, and the output of the first enhancement and suppression module is the seventh feature map.

[0095] It should be noted that the multi-scale foreground enhancement and background suppression module mentioned in the present application is the first enhancement and suppression module, which is used to obtain the enhanced feature map of the operation personnel target. This feature map can highlight the salient features of power grid operation personnel while suppressing complex background interference. The grouped feature interaction module mentioned in the present application is the second grouped feature interaction module; the lightweight extensible target tracking module mentioned in the present application is the third target tracking module.

[0096] In an embodiment of the present application, generating the second feature map according to the second data set includes:

[0097] Three paths are established to obtain three different feature maps according to the second data set;

[0098] The three paths correspond to maximum pooling operations of different sizes;

[0099] The padding operation is performed on the three different feature maps, and the sizes of the three different feature maps after the padding operation are the same.

[0100] In an optional embodiment, it is possible to design Figure 3 The first enhancement and suppression module shown, wherein:

[0101] The video frames from the image frame dataset (the second dataset) are fed into the first enhancement and suppression module to generate an enhanced feature map of the workers. This map highlights the distinctive features of power grid workers while suppressing complex background interference. The multi-scale foreground enhancement and background suppression module improves the perception of objects of different scales through multi-scale convolution. Combined with the foreground enhancement mechanism, it further enhances the features of small-scale and distant workers.

[0102] Specifically, the construction and execution process of the first enhancement and suppression module is as follows:

[0103] The video frame is input into the multi-scale foreground enhancement and background suppression module. The input video frame is convolved using a 3×3 convolution kernel, and then normalized and activated (i.e. Figure 3 The CBS operation in the above example is performed to obtain the operator target enhanced feature map F1. Then, F1 is input into the MSEA module to obtain the operator target enhanced feature map F2. Next, the CBS operation is performed on F2 again to obtain the operator target enhanced feature map F3. F3 is input into the MSEA module again to obtain the further optimized operator target enhanced feature map F4. The Conv1×1 convolution operation is then performed on F4 to extract features to obtain the operator target enhanced feature map F5. The Sigmoid operation is then performed on F5 to obtain the feature map probability distribution. ,Will The background suppression information is input into the formula for processing to generate the operator target enhancement feature map F6. F6 is concatenated with F4 in the channel dimension through the Concat operation to generate the final output operator target enhancement feature map F7.

[0104] In an embodiment of the present application, generating a second feature map according to the second data set includes:

[0105] Establish three paths and obtain three different feature maps based on the second data set;

[0106] The three paths correspond to maximum pooling operations of different sizes;

[0107] The padding operation is performed on the three different feature maps, and the sizes of the three different feature maps after the padding operation are the same.

[0108] In the embodiment of the present application, generating the second feature map according to the second data set further includes:

[0109] Perform several convolution operations on the second data set and perform upsampling operations to obtain another feature map of the same size as the three different feature maps;

[0110] Perform a splicing operation on three different feature maps and another feature map;

[0111] And perform a first feature enhancement extraction operation on the feature map after the splicing operation to obtain a second feature map.

[0112] In this embodiment of the present application, the first feature enhancement extraction operation includes:

[0113] Perform two convolution operations of different sizes on the feature map after the splicing operation, and record the result as the first convolution feature map, that is, Figure 4 T9, and the second convolutional feature map, i.e. Figure 4 Middle T10;

[0114] Generate target enhancement weights for the second convolution feature map through activation function;

[0115] The target enhancement weight is multiplied element by element with the first convolution feature map, and the element multiplication result is upsampled to obtain the second feature map, which is as follows Figure 4 The operation part after T7.

[0116] In an optional embodiment, F1 is input into the Multi-Scale Enhanced Attention (MSEA) module, i.e., the multi-scale enhanced attention submodule. The structure of the MSEA module can be as follows: Figure 4 As shown:

[0117] First, a Conv1×1 convolution operation is performed on F1 to extract the initial features of the target personnel and generate the grid operator feature map T1. At the same time, T1 is processed in three paths: In path 1, a maximum pooling operation of size 1×1 is performed on F1 (i.e. Figure 4 The MaxPool(1×1) operation in the path is used to obtain the power grid operator feature map T2; in the second path, a maximum pooling operation of size 9×9 is performed on F1 (i.e. Figure 4 MaxPool (9 × 9) operation in ), and by filling (i.e. Figure 4 The feature map size is adjusted to be consistent with T2, and the grid operator feature map T3 is obtained; in path three, a maximum pooling operation of size 13×13 is performed on F1 (i.e. Figure 4 The MaxPool(13×13) operation in

[13] is also adjusted to the same size as T2 by padding to obtain the power grid operator feature map T4.

[0118] Then, T1 is input into a convolutional module of size 3×3 (i.e. Figure 4 Conv3×3 in the image), generate the power grid operator feature map T5; extract deep features from T5 through the Conv3×3 convolution operation to generate the power grid operator feature map T6; perform upsampling operation on T6 (i.e. Figure 4 Upsample in T2 is adjusted to the same size as T2 to obtain the power grid operator characteristic map T7.

[0119] Then, T2, T3, T4, and T7 are spliced ​​along the channel dimension (i.e. Figure 4 ), and obtain the fused power grid operator feature map T8.

[0120] Next, a Conv1×1 convolution operation is performed on T8 to extract the fusion features and generate the power grid operator feature map T9; then, a Conv3×3 convolution operation is performed on T9 to obtain the power grid operator feature map T10, and the target enhancement weight is generated through the Sigmoid activation function.

[0121] Finally, the target enhancement weight is multiplied element-wise with T9 (i.e. Figure 4 * operation in the output), the grid operator feature map T11 is obtained, and an upsampling operation is performed on T11 to generate the final output operator target enhanced feature map F2.

[0122] In an optional embodiment, the background suppression information formula in the first enhancement and suppression module can be shown as follows:

[0123] ,

[0124] in, Enhance the feature map for the target. ξ( ) is the score of the feature. C is a constant used to ensure that the result is positive. is the probability distribution of the feature map.

[0125] It should be noted that the construction and operation process of the background suppression information formula is as follows: Input the background suppression information formula and use the background suppression information formula to Scoring is performed, and by setting a threshold, high-scoring areas are regarded as foregrounds and low-scoring areas as backgrounds, and then the target enhanced feature map F6 is output.

[0126] In the embodiment of the present application, the second grouping feature interaction module includes a first splitting operation;

[0127] The first splitting operation is used to split the feature map into several sub-maps;

[0128] Each subgraph is equipped with two subgraph processing paths, one of which is used to generate intra-group weights, and the other is used to generate interactive fusion feature maps.

[0129] The second grouping feature interaction module is used to output image detection frame information;

[0130] In an optional embodiment, the second grouping feature interaction module can be designed as follows Figure 5 As shown, where:

[0131] First, F7 is subjected to an upsampling operation (i.e. Figure 5 Upsample in F2 is adjusted to the same size as F2 to generate feature map F8;

[0132] Then, F2 and F8 are spliced ​​together (i.e. Figure 5 Concat in ( ) is fused in the channel dimension to obtain the feature map F9;

[0133] Next, perform a Conv 1×1 convolution operation on F9 to extract the fused preliminary features and generate the feature map F10;

[0134] Then, F10 is activated by the Sigmoid function to generate a weight map, and the weight map is multiplied element-wise with F9 (i.e. Figure 5 * operation in ), and obtain the feature map F11.

[0135] Then, split F11 according to the channel (ie Figure 5 ), is divided into three sub-graphs, namely F11_1, F11_2, and F11_3. Each sub-graph goes through the following processing flow:

[0136] The first subgraph is fed into two processing paths: in the first path, a pooling operation is first performed (i.e. Figure 5 Pool in the pool), and obtain the group interaction fusion feature map A1. Then A1 is activated by the Sigmax function to generate the group weights; in the second path, F11_1 is subjected to the Conv3×3 convolution operation to obtain the group interaction fusion feature map A2. Finally, A2 is element-wise multiplied with the group weights (i.e. Figure 5 The group interaction fusion feature map A3 is generated.

[0137] The second subgraph is fed into two processing paths respectively: In the third path, a pooling operation is first performed (i.e. Figure 5), and obtain the group interaction fusion feature map A4. Then A4 is activated by the Sigmax function to generate the group weights; in the fourth path, F11_2 is subjected to the Conv3×3 convolution operation to obtain the group interaction fusion feature map A5. Finally, A5 is element-wise multiplied with the group weights (i.e. Figure 5 The group interaction fusion feature map A6 is generated.

[0138] The third subgraph is fed into two processing paths respectively: In the fifth path, a pooling operation is first performed (i.e. Figure 5 ), and obtain the group interaction fusion feature map A7. Then A7 is activated by the Sigmax function to generate the group weights; in the sixth path, F11_3 is subjected to the Conv3×3 convolution operation to obtain the group interaction fusion feature map A8. Finally, A8 is element-wise multiplied with the group weights (i.e. Figure 5 The group interaction fusion feature map A9 is generated.

[0139] The sub-feature maps A3, A6, and A9 generated after each group processing are spliced ​​through the channel dimension (i.e. Figure 5 C operation in ), forming the fused feature map F12.

[0140] Next, a Conv1×1 convolution operation is performed on F12 to extract the final fusion features, generate the feature map F13, and generate the final weight map through the Sigmoid activation function; then, the weight map is multiplied element-by-element with F12 (i.e. Figure 5 × operation in the figure), and obtain the feature map F14. F14 is concatenated (i.e. Figure 5 The Concat in

[15] is fused with F7 to complete the feature extraction and fusion of the group feature interaction module to obtain the feature map X.

[0141] Finally, X is input into the Detection head to obtain the image detection box information, including the grid operator's bounding box coordinates and confidence score.

[0142] In an embodiment of the present application, the third target tracking module includes three different branches, namely, a first branch, a second branch and a third branch. The first branch is used to extract compact features in the image detection frame information, the second branch is used to extract local deep features in the image detection frame information, and the third branch is used to extract lightweight features in the image detection frame information.

[0143] In an optional embodiment, the third target tracking module can be designed as follows Figure 6 As shown, Figure 6 A lightweight and scalable target tracking model in

[15] , where:

[0144] The image detection box information is input into three different branches: The first branch extracts compact features through Conv1×1 convolution operation to generate feature map F15; F15 is input into CBAM (channel and spatial attention module) to generate attention feature map F18; then, F18 is input into Sigmoid activation function to generate weight feature map, and the weight map is multiplied by F15 element by element (i.e. Figure 6 * operation in the ), and obtain the feature map F22; the second branch obtains the feature map F16 through the Conv3×3 convolution operation, and performs a depth convolution operation on F16 (i.e. Figure 6 The DConv in the first branch extracts local deep features and generates feature map F19. A Conv1×1 convolution operation is performed on F19 to obtain feature map F21. The third branch directly extracts lightweight features through depthwise convolution to generate feature map F17. A Conv1×1 convolution operation is performed on F17 to obtain feature map F20.

[0145] Then, F22, F21 and F20 are concatenated in the channel dimension to generate a fused multi-scale feature map F23.

[0146] Finally, the fused feature map F22 is fed into the Kalman filter and Hungarian matching algorithm modules, which combine the target's time series information to perform target tracking. The Kalman filter module predicts the target's position, while the Hungarian matching algorithm assigns the relationship between the detection bounding box and the tracking trajectory. The final output is the target tracking result, which includes the target's trajectory ID, the detected target's position in the current frame (detection bounding box coordinates), and the target's motion trajectory.

[0147] In an optional embodiment, the Kalman filter establishes a target motion model (such as a uniform speed or uniform acceleration model) and uses the state estimate at the previous moment and the observation value at the current moment to update the estimated value of the target position, thereby achieving prediction of the future position.

[0148] The specific steps can be as follows:

[0149] Set the initial state estimate and covariance matrix during initialization;

[0150] For each frame of image, first predict the target position of the current frame based on the state of the previous frame;

[0151] Then, the actual observation results of the current frame (such as bounding box coordinates) are combined to make corrections to obtain a more accurate target position.

[0152] In an optional embodiment, the Hungarian algorithm is an optimization method for solving the assignment problem, the goal of which is to find an allocation solution that minimizes the total cost.

[0153] In this paper, each detection box is regarded as a task, and each existing track is regarded as a worker. The matching cost (such as Euclidean distance) between each detection box and each track is calculated, and then the Hungarian algorithm is applied to find the optimal matching solution.

[0154] Furthermore, based on the constructed PTRTT-Net network for real-time tracking of personnel and targets in power grid operation scenarios, the parameters of each layer were trained and updated to obtain a trained PTRTT-Net network. The basic training process is as follows: First, all neural network parameters are initialized and model-related hyperparameters are set, including the number of training rounds, batch size, weight decay coefficient, learning rate, and total number of iterations.

[0155] It should be noted that establishing a first tracking model and using the second data set as the training set and validation set of the first tracking model can effectively improve the model's ability to identify and track human targets in power grid operation scenarios. By utilizing key technologies such as multi-scale convolution, foreground enhancement mechanism, and background suppression information formula, the model can more accurately capture the salient features of the operating personnel and effectively suppress the interference of complex backgrounds. In addition, by designing modules such as the first enhancement and suppression module, the second grouping feature interaction module, and the third target tracking module, the model can achieve real-time tracking of operating personnel targets and output image detection box information containing bounding box coordinates and confidence scores, providing strong support for safety monitoring and personnel management of power grid operations.

[0156] S103 , performing real-time tracking of personnel targets in the target power grid operation scene according to the verified first tracking model.

[0157] It should be noted that after the training of the first tracking model, the real-time tracking model for personnel in power grid operation scenarios is completed, and the trained PTRTT-Net model is applied to track personnel in power grid operation scene videos in real time. During the tracking process, the human target label name, annotated box information, and trajectory data of the target over time in the video are output in real time.

[0158] In summary, the present invention proposes a method for real-time tracking of personnel targets in power grid operation scenarios, which obtains a first data set in a target power grid operation scenario, performs a first preprocessing on the first data set to obtain a second data set; establishes a first tracking model, and uses the second data set as a training set and a verification set for the first tracking model; and performs real-time tracking of personnel targets in the target power grid operation scenario based on the first tracking model after verification. The present invention solves the problems of false detection and missed detection of targets caused by background noise and illumination changes in the prior art; the method effectively highlights the key features of personnel targets through a multi-scale foreground enhancement and background suppression mechanism, while suppressing background interference, making personnel target tracking in complex power grid operation scenarios more stable and reliable; in addition, the method significantly improves the feature extraction efficiency and real-time performance in multi-target dense scenarios through grouped feature interaction optimization and lightweight network design, and reduces the computational overhead and insufficient real-time performance caused by complex network structures in the prior art, thereby achieving high-precision real-time tracking performance while ensuring lightweight.

[0159] Example 2, using Figure 2-Figure 6 The specific structure shown is used to track personnel targets in real-time in actual power grid operation scenarios, where:

[0160] The richer the variety of human targets included in a dataset, the more operational scenarios the model can adapt to. Therefore, we collected a dataset of raw videos showing real-time tracking of human targets in power grid operation scenarios. The dataset consists of 200 videos, totaling approximately 50 hours. This includes 20 hours of footage of single-person operations, 15 hours of footage of collaborative operations, 10 hours of footage of operations in dynamic environments (such as wind and passing vehicles), and 5 hours of footage of operations in extreme environments (such as nighttime, rain, and snow). The annotated "human targets" in the dataset include power grid workers in various operating situations, including solo workers, collaborative workers, and those exposed to various environmental factors (such as wind, rain, and nighttime lighting). We used the VIA (VGG Image Annotator) tool to annotate each frame of the video data and export a JSON label file containing the target's label name, annotated box information, and the target's trajectory over time within the video.

[0161] Further, the power grid operation scene personnel target real-time tracking original video dataset contains 200 videos, with a total length of about 50 hours, and about 450,000 image frames are obtained after frame-by-frame extraction. The extracted image frames are uniformly adjusted to a resolution of 1280x720 pixels. The original dataset is expanded through data enhancement methods, including: random rotation, implemented using the cv2.warpAffine() method in the OpenCV library of Python; horizontal flip, implemented using the cv2.flip() function in the OpenCV library; light change simulation, adjusted using the ImageEnhance.Brightness function of the Pillow library; adding motion blur, implemented using a custom convolution kernel in OpenCV; color jitter, simulated using the ImageEnhance.Color method. After data enhancement, the image frame dataset finally contains 1,800,000 image frames.

[0162] Further, the image frame dataset is divided according to a ratio of 8:1:1, resulting in 1,440,000 frames for the training set, 180,000 frames for the validation set, and 180,000 frames for the test set, ensuring balanced distribution of training, validation, and testing among video samples under different environments and scenarios.

[0163] Further, the image frame with a resolution of 1280x720 pixels is input into the multi-scale foreground enhancement and background suppression module, and the specific process is as follows: a 3x3 convolution kernel is used to perform convolution operation on the input image frame, and then normalization and activation function processing are performed (i.e. Figure 3The CBS operation in

[15] is performed to generate an enhanced operator target feature map F1 with a size of 1280×720 pixels and 32 channels. Feature map F1 is input into the MSEA module to obtain a fused enhanced operator target feature map F2 with a size of 1280×720 pixels and 96 channels. The CBS operation is performed again on the enhanced operator target feature map F2 to generate an enhanced operator target feature map F3 with a size of 1280×720 pixels and 96 channels. F3 is input into the MSEA module and the multi-scale feature extraction process is repeated to obtain an optimized enhanced operator target feature map F4 with a size of 1280×720 pixels and 96 channels. The optimized feature map F4 is convolved with a 1×1 convolution kernel to extract features and generate an enhanced operator target feature map F5 with a size of 1280×720 pixels and 32 channels. Process F5 using the Sigmoid activation function to obtain the probability distribution p(x) of the feature map. Process p(x) according to the background suppression information formula to generate the enhanced worker target feature map F6, which is 1280×720 pixels and has 32 channels. Use the concatenation operation (Concat) to fuse feature maps F6 and F4 along the channel dimension to generate the final output enhanced worker target feature map F7, which is 1280×720 pixels and has 128 channels.

[0164] Furthermore, in the MSEA module, a 1×1 convolution operation is performed on the input feature map F1 to extract initial features, generating a grid worker feature map T1 with a size of 1280×720 pixels and 32 channels. T1 is processed in three paths: In the first path, a 1×1 pooling kernel is used to generate a grid worker feature map T2 with a size of 1280×720 pixels and 32 channels. In the second path, a 9×9 pooling kernel is used and the feature map is resized to match T2 through padding, generating a grid worker feature map T3 with a size of 1280×720 pixels and 32 channels. In the third path, a 13×13 pooling kernel is used and the feature map is resized to match T2 through padding, generating a grid worker feature map T4 with a size of 1280×720 pixels and 32 channels. T1 is fed into a 3×3 convolutional module to generate a grid worker feature map T5, with a size of 1280×720 pixels and 32 channels. T5 is subjected to another 3×3 convolution operation to extract deep features, generating a grid worker feature map T6, with a size of 1280×720 pixels and 32 channels. T6 is upsampled to the same size as T2, T3, and T4, generating a grid worker feature map T7, with a size of 1280×720 pixels and 32 channels. T2, T3, T4, and T7 are then concatenated along the channel dimension to generate a fused grid worker feature map T8, with a size of 1280×720 pixels and 128 channels. T8 is subjected to a 1×1 convolution operation to extract fused features, generating a grid worker feature map T9, with a size of 1280×720 pixels and 64 channels. T9 is fed into a 3×3 convolutional module to extract deep fusion features, generating a power grid operator feature map T10 with a size of 1280×720 pixels and 64 channels. T10 is processed using a sigmoid activation function to generate a weight map with a size of 1280×720 pixels and 64 channels. The weights are element-wise multiplied by T9 to generate a power grid operator feature map T11 with a size of 1280×720 pixels and 64 channels. Finally, a 1×1 convolution operation is performed on T11 to fuse the information, generating the final output feature map F2 with a size of 1280×720 pixels and 96 channels.

[0165] Further, in the group feature interaction module, first, the input feature map F7 (size 1280x720 pixels, 128 channels) is adjusted to the same size as F2 (size 1280x720 pixels, 96 channels) through upsampling operation, to generate a feature map F8, which has a size of 1280x720 pixels, 128 channels. Then, F2 and F8 are fused in the channel dimension through splicing operation to obtain a feature map F9, which has a size of 1280x720 pixels, 224 channels. Next, F9 is input into a convolution module with a size of 1x1 to extract the fused preliminary features and generate a feature map F10, which has a size of 1280x720 pixels, 32 channels. Subsequently, F10 is processed through a Sigmoid activation function to generate a weight map, which has a size of 1280x720 pixels, 3 channels. The weight map is multiplied element by element with F9 to obtain a feature map F11, which has a size of 1280x720 pixels, 224 channels. Then, F11 is split into 3 subgraphs: F11_1, F11_2, F11_3 according to the channel, each of which has a size of 1280x720 pixels, 74 channels. Each subgraph is processed through the following processing flow: for the first subgraph F11_1, it is input into two processing paths: in the first path, first, a pooling operation is performed to obtain a group interaction fusion feature map A1, which has a size of 1280x720 pixels, 74 channels; then A1 is processed through a Sigmoid activation function to generate intra-group weights, which have a size of 1280x720 pixels, 74 channels. In the second path, F11_1 is processed through a 3x3 convolution operation to obtain a group interaction fusion feature map A2, which has a size of 1280x720 pixels, 74 channels; finally, A2 is multiplied element by element with the intra-group weights to generate a group interaction fusion feature map A3, which has a size of 1280x720 pixels, 74 channels. For the second subgraph F11_2, it is also input into two processing paths: in the third path, first, a pooling operation is performed to obtain a group interaction fusion feature map A4, which has a size of 1280x720 pixels, 74 channels; then A4 is processed through a Sigmoid activation function to generate intra-group weights, which have a size of 1280x720 pixels, 74 channels. In the fourth path, F11_2 is processed through a 3x3 convolution operation to obtain a group interaction fusion feature map A5, which has a size of 1280x720 pixels, 74 channels; finally, A5 is multiplied element by element with the intra-group weights to generate a group interaction fusion feature map A6, which has a size of 1280x720 pixels, 74 channels.The third sub-image, F11_3, is similarly fed into two processing paths. In the fifth path, pooling is first performed to generate a group-interaction fused feature map, A7, with a size of 1280×720 pixels and 74 channels. A7 is then activated with a sigmoid function to generate intra-group weights, also with a size of 1280×720 pixels and 74 channels. In the sixth path, a 3×3 convolution operation is performed on F11_3 to generate a group-interaction fused feature map, A8, with a size of 1280×720 pixels and 74 channels. Finally, A8 is element-wise multiplied with the intra-group weights to generate a group-interaction fused feature map, A9, with a size of 1280×720 pixels and 74 channels. The sub-feature maps A3, A6, and A9 generated by each sub-image processing are concatenated along the channel dimension to form the fused feature map, F12, with a size of 1280×720 pixels and 222 channels. Next, F12 is fed into a 1×1 convolutional module to extract the final fused features, generating a feature map F13 with a size of 1280×720 pixels and 32 channels. A sigmoid activation function is then used to generate a final weight map with a size of 1280×720 pixels and 32 channels. Subsequently, the weight map is element-wise multiplied with F12 to generate a feature map F14 with a size of 1280×720 pixels and 222 channels. Finally, F14 is concatenated with F7 to generate the final feature map X with a size of 1280×720 pixels and 350 channels. This feature map X is fed into the detection head to obtain the detection box information for each frame.

[0166] Further, in the lightweight scalable object tracking module, the bounding box information (size of 1280x720 pixels, 350 channel numbers) of each frame image is input to three different branches for processing. The first branch extracts compact features through a convolution kernel with a size of 1x1 (i.e. Conv 1x1 in the figure), to generate a feature map F15 with a size of 1280x720 pixels, 32 channel numbers. F15 is input to CBAM (Channel and Spatial Attention Module) to generate an attention feature map F18 with a size of 1280x720 pixels, 32 channel numbers. Then, F18 is input to a Sigmoid activation function to generate a weight feature map with a size of 1280x720 pixels, 32 channel numbers. The weight map is multiplied element by element with F15 to obtain a feature map F22 with a size of 1280x720 pixels, 32 channels. The second branch obtains a feature map F16 with a size of 1280x720 pixels, 64 channel numbers through a convolution kernel with a size of 3x3 (i.e. Conv 3x3 in the figure). F16 is subjected to a deep convolution operation (i.e. DConv in the figure) to extract local deep features, to generate a feature map F19 with a size of 1280x720 pixels, 64 channel numbers. F19 is subjected to a 1x1 convolution operation to obtain a feature map F21 with a size of 1280x720 pixels, 32 channel numbers. The third branch directly extracts lightweight features through a deep convolution operation (i.e. DConv in the figure) to generate a feature map F17 with a size of 1280x720 pixels, 32 channel numbers. F17 is subjected to a 1x1 convolution operation to obtain a feature map F20 with a size of 1280x720 pixels, 32 channel numbers.

[0167] Further, F22, F21 and F20 are spliced in the channel dimension to generate a fused multi-scale feature map F23 with a size of 1280x720 pixels, 96 channel numbers. Finally, the fused feature map F23 is input to a Kalman filtering and Hungarian matching module. The Kalman filtering module is responsible for predicting the position state of the target, to generate a target prediction position with a size of 1280x720 pixels, 32 channel numbers; the Hungarian matching algorithm is responsible for assigning the association relationship between the detection box and the tracking trajectory, to finally output the tracking result of the target, including the trajectory ID of the target, the position of the current frame detection target (detection box coordinates) and the motion trajectory of the target.

[0168] Furthermore, a real-time tracking network for personnel in power grid operation scenarios, PTRTT-Net, was trained over 120 epochs, with a weight decay coefficient of 0.005, an initial learning rate of 0.0001, and a batch size of 16. The number of iterations per round was 250 for the first 60 epochs and 350 for the next 60 epochs, for a total of 36,000 iterations. During each training round, random data augmentation techniques (such as random cropping, rotation, horizontal flipping, illumination change simulation, and motion blur) were used to further improve the model's generalization capabilities. The input video frames were uniformly resized to 1280×720 pixels to accommodate the network input requirements. After parameter initialization, the processed training and validation datasets for personnel in power grid operation scenarios were divided into multiple batches, and one batch of data was fed into PTRTT-Net for training. PGT-Net calculated the training loss for each batch and dynamically updated the model parameters based on the loss using the backpropagation algorithm.

[0169] After all batches of data from the entire training set have been trained, the validation phase begins. During this phase, the validation set data is fed into PTRTT-Net batch by batch, and the validation loss (batch_loss) for each batch is calculated. Based on the training loss (loss) and validation loss (batch_loss) of each round of training, PGT-Net dynamically adjusts the learning rate and optimizes parameters to improve training efficiency and model performance.

[0170] In an optional embodiment, the training loss value (loss) generally refers to a measure of the difference between the model output and the true label, and can use a cross entropy loss function or mean square error, etc. For target tracking tasks, the loss function may also include factors such as position offset and size change.

[0171] In an optional embodiment, the validation loss (batch_loss) refers to the average loss calculated by the model on the validation set during the validation phase. The validation loss (batch_loss) is similar to the training loss, but is calculated based on the validation set data and is used to evaluate the model's performance on unseen data. During each validation phase, a batch of validation set data is fed into the model, and the average loss for all samples in that batch is calculated.

[0172] It should be noted that the entire training process continues through multiple iterations until the validation loss (batch_loss) converges, indicating that the model has achieved good real-time tracking capabilities and training is complete. Ultimately, the trained PTRTT-Net is capable of high-precision real-time tracking of personnel targets in power grid operation scenarios, maintaining stable and efficient tracking even in complex environments (such as multi-person collaboration, background interference, and extreme weather).

[0173] Example 3: This embodiment also provides a real-time tracking system for personnel targets in power grid operation scenarios, including:

[0174] A data acquisition and processing module, configured to acquire a first data set in a target power grid operation scenario, and perform a first preprocessing on the first data set to obtain a second data set;

[0175] A model building module, configured to build a first tracking model and use the second data set as a training set and a validation set for the first tracking model;

[0176] The first tracking model includes a first enhancement and suppression module, a second grouping feature interaction module and a third target tracking module;

[0177] The first enhancement and suppression module includes a multi-scale enhancement attention submodule, and the first enhancement and suppression module inputs the second feature map and the seventh feature map into the second grouping feature interaction module;

[0178] The multi-scale enhanced attention submodule is used to generate a second feature map based on the second data set, and the output of the first enhancement and suppression module is the seventh feature map;

[0179] The second grouping feature interaction module is used to output image detection frame information;

[0180] The tracking module is used to perform real-time tracking of personnel targets in the target power grid operation scene based on the first tracking model after verification.

[0181] The above-mentioned unit modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of the above-mentioned modules.

[0182] This embodiment also provides a computer device, which may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7As shown. The computer device includes a processor, memory, communication interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for real-time tracking of personnel targets in power grid operation scenarios is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0183] This embodiment further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the following steps are implemented:

[0184] Acquire a first data set in a target power grid operation scenario, and perform a first preprocessing on the first data set to obtain a second data set;

[0185] Establish a first tracking model, and use the second data set as a training set and a validation set for the first tracking model;

[0186] The first tracking model includes a first enhancement and suppression module, a second grouping feature interaction module and a third target tracking module;

[0187] The first enhancement and suppression module includes a multi-scale enhancement attention submodule, and the first enhancement and suppression module inputs the second feature map and the seventh feature map into the second grouping feature interaction module;

[0188] The multi-scale enhanced attention submodule is used to generate a second feature map based on the second data set, and the output of the first enhancement and suppression module is the seventh feature map;

[0189] The second grouping feature interaction module is used to output image detection frame information;

[0190] Real-time tracking of personnel targets in target power grid operation scenarios is performed based on the first tracking model after verification.

[0191] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

[0192] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application may be implemented using various computer languages.

[0193] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0194] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0195] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0196] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0197] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for real-time tracking of personnel targets in power grid operation scenarios, characterized in that: include: Acquire a first data set in a target power grid operation scenario, and perform a first preprocessing on the first data set to obtain a second data set; Establishing a first tracking model, and using the second data set as a training set and a validation set for the first tracking model; The first tracking model includes a first enhancement and suppression module, a second grouping feature interaction module and a third target tracking module; The first enhancement and suppression module includes a multi-scale enhancement attention submodule, and the first enhancement and suppression module inputs the second feature map and the seventh feature map into the second grouping feature interaction module; Input the video frame into the multi-scale foreground enhancement and background suppression module; Use a 3×3 convolution kernel to perform a convolution operation on the input video frame, then perform normalization and activation function processing to obtain the operator target enhanced feature map F1; Input F1 into the multi-scale enhanced attention submodule to obtain the second feature map F2 of the operator target enhancement; Normalize and activate the second feature map F2 to obtain the operator target enhanced feature map F3, and input F3 into the multi-scale enhanced attention submodule again to obtain a further optimized operator target enhanced feature map F4; Then, a Conv1×1 convolution operation is performed on F4 to extract features, obtaining the operator target enhanced feature map F5, and a Sigmoid activation function operation is performed on F5 to obtain the feature map probability distribution p(x); Input p(x) into the background suppression information formula to generate the operator target enhancement feature map F6, and then splice F6 with F4 in the channel dimension to generate the seventh feature map F7 of the operator target enhancement. The multi-scale enhanced attention submodule is used to generate a second feature map based on the second data set, and the output of the first enhancement and suppression module is a seventh feature map; The multi-scale enhanced attention submodule includes: Perform a Conv1×1 convolution operation on F1 to extract the initial features of the target operator and generate a grid operator feature map T1. T1 is divided into three paths for processing: In path 1, a maximum pooling operation of size 1×1 is performed on F1 to obtain the grid operator feature map T2; In the second path, a maximum pooling operation of 9×9 is performed on F1, and the feature map size is adjusted by padding to make it consistent with T2, thus obtaining the power grid operator feature map T3; In path three, a maximum pooling operation of size 13×13 is performed on F1, and it is also adjusted to the same size as T2 by padding to obtain the power grid operator feature map T4; Input T1 into a 3×3 convolution module to generate a grid operator feature map T5; The deep features of T5 are extracted through Conv3×3 convolution operation to generate the power grid operator feature map T6; Perform upsampling on T6 and adjust it to the same size as T2 to obtain the grid operator feature map T7; T2, T3, T4, and T7 are spliced ​​along the channel dimension to obtain the fused power grid operator characteristic map T8; Perform a Conv1×1 convolution operation on T8 to extract fusion features and generate a power grid operator feature map T9; Perform a Conv3×3 convolution operation on T9 to obtain the power grid operator feature map T10, and generate the target enhancement weight through the Sigmoid activation function; Multiply the target enhancement weight by T9 element by element to obtain the power grid operator feature map T11, perform an upsampling operation on T11, and generate the final output operator target enhanced second feature map F2; The second grouping feature interaction module is used to output image detection frame information; Real-time tracking of personnel targets in target power grid operation scenarios is performed based on the first tracking model after verification.

2. The method for real-time tracking of personnel targets in power grid operation scenarios according to claim 1, characterized in that: Generating a second feature map according to the second data set includes: Establishing three paths to obtain three different feature maps based on the second data set; The three paths correspond to maximum pooling operations of different sizes; A padding operation is performed on the three different feature maps, and sizes of the three different feature maps after the padding operation are the same.

3. The method for real-time tracking of personnel targets in power grid operation scenarios according to claim 2, characterized in that: Generating a second feature map according to the second data set further includes: Performing several convolution operations on the second data set and performing an upsampling operation on the second data set to obtain another feature map of the same size as the three different feature maps; Performing a splicing operation on the three different feature maps and another feature map; A first feature enhancement extraction operation is performed on the feature map after the splicing operation to obtain a second feature map.

4. The method for real-time tracking of personnel targets in power grid operation scenarios according to claim 3, characterized in that: The first feature enhancement extraction operation includes: Performing two convolution operations of different sizes on the feature map after the splicing operation, and recording the results as a first convolution feature map and a second convolution feature map; Generating a target enhancement weight for the second convolutional feature map through an activation function; The target enhancement weight is multiplied by the first convolution feature map element by element, and the element-wise multiplication result is upsampled to obtain a second feature map.

5. The method for real-time tracking of personnel targets in power grid operation scenarios according to claim 4, characterized in that: The performing a first preprocessing on the first data set includes: Annotate the first dataset, use a video annotation tool to select the target position of the person in the video, and annotate the label; The frame selection operation is to determine the horizontal and vertical coordinates of the upper left corner and the lower right corner of the frame in each frame; The label annotation includes the target person's label name, annotation box information, and the target's trajectory changing over time in the video; The first data set is a video data set.

6. The method for real-time tracking of personnel targets in power grid operation scenarios according to claim 5, characterized in that: The second grouping feature interaction module includes a first splitting operation; The first splitting operation is used to split the feature map into a plurality of sub-maps; Each subgraph is equipped with two subgraph processing paths, one of which is used to generate intra-group weights, and the other is used to generate interactive fusion feature maps.

7. The method for real-time tracking of personnel targets in power grid operation scenarios according to claim 6, characterized in that: The third target tracking module includes three different branches, namely a first branch, a second branch and a third branch. The first branch is used to extract compact features in the image detection frame information, the second branch is used to extract local deep features in the image detection frame information, and the third branch is used to extract lightweight features in the image detection frame information.

8. A real-time tracking system for personnel targets in power grid operation scenarios, applying the method according to any one of claims 1 to 7, characterized in that: include: a data acquisition and processing module, configured to acquire a first data set in a target power grid operation scenario, and perform a first preprocessing on the first data set to obtain a second data set; a model building module, configured to build a first tracking model, and use the second data set as a training set and a validation set for the first tracking model; The first tracking model includes a first enhancement and suppression module, a second grouping feature interaction module and a third target tracking module; The first enhancement and suppression module includes a multi-scale enhancement attention submodule, and the first enhancement and suppression module inputs the second feature map and the seventh feature map into the second grouping feature interaction module; The multi-scale enhanced attention submodule is used to generate a second feature map based on the second data set, and the output of the first enhancement and suppression module is a seventh feature map; The second grouping feature interaction module is used to output image detection frame information; The tracking module is used to perform real-time tracking of personnel targets in the target power grid operation scene based on the first tracking model after verification.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Target behavior detection model construction method, system and device and storage medium

    CN119229518A

  • Improved Yolov10-oriented target detection fusion method

    CN119723272A