HOI-based power distribution network live working behavior prediction method, apparatus and device

Through multi-scale feature extraction and optimization anchor point technology based on HOI, the detection problem of live operation behavior in the distribution network under complex background is solved, and higher-precision safety risk identification is achieved.

CN120375475APending Publication Date: 2025-07-25XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510516090.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to extract key live operation features from images containing complex background information, making it difficult to effectively identify security risks in multi-factor interaction.

Method used

Using a HOI-based method, fine-grained anchor points are optimized through multi-scale feature extraction, deformable Transformer encoder, pyramid space attention module and task-aware decoupling module, combined with topological alignment, accurate prediction of live job behavior is achieved.

Benefits of technology

It improves the accuracy of live-operated operation behavior detection in the distribution network, effectively suppresses background noise, enhances the identification ability of multi-scale spatial relationships, and improves the accuracy of safety risk identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375475A_ABST
    Figure CN120375475A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network hot-line work behavior prediction method, device and equipment based on HOI, and the method comprises the steps: carrying out the multi-scale feature extraction of obtained target hot-line work image data in a power distribution network, and obtaining the multi-scale hot-line work features; semantic coding is carried out on the multi-scale hot-line work features to obtain coded hot-line work features; embedding an initial anchor point into a decoder in the trained hot-line work behavior recognition model, and sequentially iteratively optimizing the fine-grained anchor point through a multi-scale sampling module, a pyramid space attention module and a task perception decoupling module in the trained hot-line work behavior recognition model to obtain an optimized fine-grained anchor point; carrying out topological alignment on the optimized fine-grained anchor points and the encoded hot-line work features to obtain target hot-line work features; and performing behavior prediction on the target hot-line work image data based on the target hot-line work characteristics to obtain a behavior prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power vision detection, and particularly to a method, device and equipment for predicting live working behaviors of a distribution network based on HOI. Background Art

[0002] At present, the intelligent detection of safety risks in the production and operation links of distribution networks at home and abroad focuses on aspects such as the correct wearing of safety protection tools, the fault diagnosis of distribution equipment, and the identification of staff's illegal behaviors. By means of high-performance image perception devices such as fixed video monitoring, robots, safety control balls, law enforcement instruments, intelligent helmets, and infrared cameras, image data related to power production safety is collected, and then computer technology means such as object detection, object tracking, image segmentation, and image classification are used to analyze potential safety hazards, so as to realize the intelligent early warning function for safety risks and hidden dangers in the distribution network.

[0003] At present, the research on safety risk identification in distribution network scenarios mostly focuses on the detection of single risk factors. However, the actual operation scenario of the distribution network is a complex process involving the interaction of multiple elements such as personnel, tools, and equipment machinery. With the continuous expansion of the scale of the distribution network, the types of safety risks are increasing day by day, and the traditional single risk detection method is difficult to meet the needs of detecting interactive safety risks in complex distribution network operation scenarios.

[0004] To solve this problem, more and more detection algorithms for interaction relationships have been successively applied to the distribution network operation scenario. However, as an important problem in computer vision, Human Object Interaction (HOI) needs to locate the human-object pair and identify the interaction relationship between them. Compared with a single object instance, the HOI instance has a greater span in space, scale, and task, making its detection more vulnerable to interference from noisy backgrounds.

[0005] Therefore, how to reduce the interference of the noisy background on HOI detection and how to extract key features from images containing complex background information is a technical problem that needs to be solved urgently. Summary of the Invention

[0006] In order to solve the above problems existing in the prior art, the present invention provides a method, device and equipment for predicting live working behaviors of a distribution network based on HOI, so as to solve the problem of difficult extraction of key operation features from images containing complex background information in the prior art.

[0007] To achieve the above object, the technical solution of the embodiment of the present invention is as follows:

[0008] In the first aspect, the present invention provides a method for predicting live working behaviors of a distribution network based on HOI, and the method includes:

[0009] Obtain target live working image data in the distribution network;

[0010] Perform multi-scale feature extraction on the target live working image data to obtain multi-scale live working features;

[0011] Based on the deformable Transformer encoder in the trained live working behavior recognition model, perform semantic encoding on the multi-scale live working features to obtain encoded live working features;

[0012] Embed an initial anchor point in the decoder of the trained live working behavior recognition model, and sequentially iterate and optimize the fine-grained anchor point through the multi-scale sampling module, pyramid spatial attention module, and task-aware decoupling module in the trained live working behavior recognition model to obtain an optimized fine-grained anchor point;

[0013] Perform topological alignment on the optimized fine-grained anchor point and the encoded live working features to obtain target live working features;

[0014] Based on the target live working features, perform behavior prediction on the target live working image data to obtain a behavior prediction result.

[0015] In a second aspect, the present invention provides a device for predicting live working behavior in a distribution network based on HOI, and the device includes:

[0016] An acquisition module, configured to obtain target live working image data in the distribution network;

[0017] An extraction module, configured to perform multi-scale feature extraction on the target live working image data to obtain multi-scale live working features;

[0018] An encoding module, configured to perform semantic encoding on the multi-scale live working features based on the deformable Transformer encoder in the trained live working behavior recognition model to obtain encoded live working features;

[0019] An optimization module, configured to embed an initial anchor point in the decoder of the trained live working behavior recognition model, and sequentially iterate and optimize the fine-grained anchor point through the multi-scale sampling module, pyramid spatial attention module, and task-aware decoupling module in the trained live working behavior recognition model to obtain an optimized fine-grained anchor point;

[0020] An alignment module, configured to perform topological alignment on the optimized fine-grained anchor point and the encoded live working features to obtain target live working features;

[0021] A prediction module, configured to perform behavior prediction on the target live working image data based on the target live working characteristics, and obtain a behavior prediction result.

[0022] In some embodiments, the acquisition module is further configured to acquire initial live working image data in the distribution network; remove images with severe occlusion and images that do not contain HOI embodiments from the initial live working image data to obtain the removed live working image data; label the removed live working image data to obtain live working image data with labeled HOI information; perform data augmentation on the live working image data with labeled HOI information to obtain the target live working image data.

[0023] In some embodiments, the extraction module is further configured to perform multi-scale feature extraction on the target live working image data based on the multi-scale feature extraction layer in the trained live working behavior recognition model to obtain the multi-scale live working characteristics; wherein, the multi-scale feature extraction layer includes a hierarchical backbone network and a deformable transformer encoder, and the formula is as follows: In the formula, F e ncoder represents the encoder, F f latten represents the flattening operation, φ(·) represents the backbone network, p is the position encoding, s is the spatial shape of the multi-scale feature, r represents the effective ratio, l represents the level index corresponding to the multi-scale feature, represents a real matrix space with N rows and Cd columns, N is the number of features, and Cd is the dimension of the feature.

[0024] In some embodiments, the optimization module is further configured to, during the decoding process, establish a spatial attention region to enable the decoder to perform selective feature aggregation on the target region, and map the optimized semantic instance embedding vector to the HOI embedding; through the multi-scale sampling module, the pyramid spatial attention module, and the task-aware decoupling module in the trained distribution network live working left-right behavior recognition model, iteratively optimize the fine-grained anchor points in sequence to obtain the optimized fine-grained anchor points.

[0025] In some embodiments, the HOI-based distribution network live working behavior prediction method is applied to a trained live working behavior recognition model, and the trained live working behavior recognition model is trained through the following steps: inputting sample live working image data into an initial live working behavior recognition model;

[0026] Through the feature extraction layer of the initial live working behavior recognition model, multi-scale feature extraction is performed on the sample live working image data to obtain sample multi-scale live working features; through the deformable Transformer encoding layer of the initial live working behavior recognition model, semantic encoding is performed on the sample multi-scale live working features to obtain the encoded sample live working features; initial sample anchors are embedded in the decoder of the initial live working behavior recognition model; through the multi-scale sampling module, pyramid spatial attention module and task-aware decoupling module of the initial live working behavior recognition model, the sample fine-grained anchors are iteratively optimized in sequence to obtain the optimized sample fine-grained anchors; the optimized sample fine-grained anchors and the encoded sample live working features are topologically aligned to obtain the target sample live working features; through the behavior prediction layer of the initial live working behavior recognition model, based on the target sample live working features, behavior prediction is performed on the target sample live working image data to obtain the sample behavior prediction result; the sample behavior prediction result is input into a preset loss model to obtain a loss result; based on the loss result, the model parameters in the initial live working behavior recognition model are corrected to obtain a trained live working behavior recognition model.

[0027] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory for storing executable instructions; a processor for implementing the above-mentioned method for predicting live working behavior of a distribution network based on HOI when executing the executable instructions stored in the memory.

[0028] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing executable instructions for causing a processor to implement the above-mentioned method for predicting live working behavior of a distribution network based on HOI when executing the executable instructions.

[0029] The method for predicting the live working behavior of a distribution network based on HOI provided by the present invention first performs multi-scale feature extraction on the obtained target live working image data in the distribution network to obtain multi-scale live working features; performs semantic encoding on the multi-scale live working features to obtain the encoded live working features; embeds initial anchor points in the decoder of the trained live working behavior recognition model, and sequentially iteratively optimizes the fine-grained anchor points through the multi-scale sampling module, the pyramid spatial attention module, and the task-aware decoupling module in the trained live working behavior recognition model to obtain the optimized fine-grained anchor points; aligns the optimized fine-grained anchor points and the encoded live working features topologically to obtain the target live working features; and performs behavior prediction on the target live working image data based on the target live working features to obtain the behavior prediction result. In this way, the present invention proposes an end-to-end distribution network safety risk recognition model based on Transformer, inserts a lightweight attention module after the multi-scale features output by its backbone network, which can suppress background noise and highlight key regions such as the human body and tools; and improves the encoder and decoder of Transformer, optimizes the computational efficiency of the attention mechanism, and proposes a pyramid spatial attention module, a task-aware decoupling fusion module, and a HOI embedding module, which perform spatial weight allocation on feature maps of different scales, enhance the multi-scale spatial relationship, design independent attention branches for interactive actions and object detection respectively, effectively improve the detection accuracy of predicting the live working behavior of the distribution network, and are verified in experiments. Description of the Drawings

[0030] Figure 1 is a schematic structural diagram of a system for predicting the live working behavior of a distribution network based on HOI provided by an embodiment of the present invention;

[0031] Figure 2 is a schematic flow diagram of a method for predicting the live working behavior of a distribution network based on HOI provided by an embodiment of the present invention;

[0032] Figure 3 is a schematic flow diagram of a method for predicting the live working behavior of a distribution network based on HOI provided by an embodiment of the present invention;

[0033] Figure 4 is a schematic structural diagram of the composition of a device for predicting the live working behavior of a distribution network based on HOI provided by an embodiment of the present invention;

[0034] Figure 5 is a schematic structural diagram of the composition of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0035] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0036] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meaning as commonly understood by those skilled in the technical field to which the embodiments of the present invention belong. The terms used in the embodiments of the present invention are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.

[0037] The following describes the exemplary application of the HOI-based power distribution network live working behavior prediction device according to the embodiments of the present invention. The HOI-based power distribution network live working behavior prediction device provided by the embodiments of the present invention can be implemented as a terminal or a server. In one implementation, the HOI-based power distribution network live working behavior prediction device provided by the embodiments of the present invention can be implemented as various types of terminals such as laptops, tablets, desktop computers, mobile devices, etc.; in another implementation, the HOI-based power distribution network live working behavior prediction device provided by the embodiments of the present invention can also be implemented as a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs, Content Delivery Networks), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present invention. Below, the exemplary application when the HOI-based power distribution network live working behavior prediction device is a server will be described.

[0038] See Figure 1 , Figure 1It is a schematic structural diagram of a power distribution network live working behavior prediction system 10 based on HOI provided by an embodiment of the present invention. An embodiment of the present invention can provide a power distribution network live working behavior prediction platform based on HOI, and the power distribution network live working behavior prediction platform based on HOI can be implemented as a power distribution network live working behavior prediction application based on HOI. The power distribution network live working behavior prediction system 10 provided by an embodiment of the present invention includes a terminal 110, a network 120, and a server 130. Among them, the server 130 is a server of a power distribution network live working behavior prediction application based on HOI. The server 130 can constitute the power distribution network live working behavior prediction device of an embodiment of the present invention. The terminal 110 is connected to the server 130 through the network 120, and the network 120 can be a wide area network, a local area network, or a combination of the two.

[0039] In some embodiments, please refer to Figure 1 , when predicting the behavior of live working in the power distribution network, the terminal 110 sends the live working behavior prediction task to the server 130 through the network 120. The server 130 responds to the live working behavior prediction task sent by the terminal 110, and acquires the target live working image data in the power distribution network; performs multi-scale feature extraction on the target live working image data to obtain multi-scale live working features; based on the deformable Transformer encoder in the trained live working behavior recognition model, performs semantic encoding on the multi-scale live working features to obtain encoded live working features; embeds an initial anchor point in the decoder of the trained live working behavior recognition model, and sequentially iteratively optimizes the fine-grained anchor point through the multi-scale sampling module, the pyramid spatial attention module, and the task-aware decoupling module in the trained live working behavior recognition model to obtain an optimized fine-grained anchor point; aligns the optimized fine-grained anchor point and the encoded live working features topologically to obtain target live working features; based on the target live working features, performs behavior prediction on the target live working image data to obtain a behavior prediction result. After obtaining the behavior prediction result, the server 130 sends the behavior prediction result to the terminal 110 through the network 120.

[0040] An embodiment of the present invention provides a power distribution network live working behavior prediction method based on HOI. Refer to Figure 2 , Figure 2 is a schematic flowchart of a power distribution network live working behavior prediction method provided by an embodiment of the present invention, and will be described in conjunction with Figure 2 the steps shown.

[0041] Step S210, acquire the target live working image data in the power distribution network.

[0042] In some embodiments, the distribution network is a specific environment where live working to be detected occurs and is also the source scenario for acquiring image data.

[0043] In some embodiments, the user can collect initial live working image data to be detected through an image acquisition component and perform a series of data processing on the initial live working image data, so as to obtain target live working image data with higher accuracy and precision. Here, the image acquisition component can be a binocular depth camera, an image sensor, or other components or devices with the function of acquiring image information. The embodiments of the present invention do not make specific limitations in this regard.

[0044] In some embodiments, the target live working image data refers to a data set in graphical form obtained through an image acquisition component for the specific target scenario of live working in the distribution network. These data sets undergo a series of data processing, such as screening, denoising, correction, etc.

[0045] In practical examples, the components in the target live working image data not only include the action information, posture information of the operator, and the relative position relationship information with the live equipment, but also include relevant information about the surrounding environment, such as the appearance size information of the equipment and the spatial information of the surrounding infrastructure. In practical applications, these target live working image data can be used to detect potential safety hazards in the live working process in real time, such as detecting whether the operator violates the operating procedures and whether the equipment is in an abnormal state, so as to improve the safety and efficiency of live working in the distribution network.

[0046] Step S220: Extract multi-scale features from the target live working image data to obtain multi-scale live working features.

[0047] In some embodiments, multi-scale feature extraction is a method of extracting feature information at different scales from the target live working image data. Its core idea is to capture features of different sizes and different levels of detail in the image by using filters of different sizes or processing the image at different resolutions.

[0048] In some embodiments, the multi-scale live working features are a set of features with different scale information obtained from the target live working image data through a multi-scale feature extraction method. These features can describe the characteristics of the live working scenario at different levels and details, and are a representation of the content of the live working image. For example, after multi-scale feature extraction, the following multi-scale live working features may be obtained. At a small scale, texture features on the gloves of the operator and shape features of the tool tips may be extracted, which help to accurately identify the working tools and hand movements; at a medium scale, the body posture features of the operator can be obtained, such as the extension angle of the arm and the bending degree of the body, which can be used to determine whether the operator's operation complies with the specifications; at a large scale, the overall structural features of the working scenario can be obtained, such as the layout of the energized equipment and the environmental features of the working area, which helps to grasp the entire live working scenario and background information macroscopically.

[0049] Step S230: Based on the deformable Transformer encoder in the trained live working behavior recognition model, perform semantic encoding on the multi-scale live working features to obtain the encoded live working features.

[0050] In practical embodiments, in the live working behavior recognition model, the deformable Transformer encoder can improve the processing efficiency and accuracy of the model for multi-scale live working features. Due to the complexity of the live working scenario, which contains a large amount of redundant information, the deformable Transformer encoder can help the model quickly locate the key features, thereby improving the performance of the model for recognizing live working behaviors.

[0051] In some embodiments, semantic encoding refers to the process of converting the input multi-scale live working features into a representation form with semantic meanings. The encoded live working features are the result obtained after the multi-scale live working features are semantically encoded by the deformable Transformer encoder.

[0052] Step S240: Embed the initial anchor points in the decoder of the trained live working behavior recognition model, and sequentially iterate and optimize the fine-grained anchor points through the multi-scale sampling module, pyramid spatial attention module, and task-aware decoupling module in the trained live working behavior recognition model to obtain the optimized fine-grained anchor points.

[0053] In some embodiments, initial anchor points are embedded in the decoder of the trained live working behavior recognition model. Here, the initial anchor points can be understood as a kind of preset reference points. Their role is similar to pre - marking some possible positions on a map, providing a basis for subsequent more accurate positioning and analysis. In the context of live working behavior recognition, these initial anchor points initially determine the regions where key features of live working behavior may exist for the recognition model. For example, in a live working image, the initial anchor points may mark the approximate position regions where the operator's body, tools, or energized equipment may appear.

[0054] In the present invention, the initial anchor points are preliminarily optimized through a multi - scale sampling module to obtain finer - grained anchor points. Since the live working scenarios are complex, including various scale information from small details (such as the operating part of a tool) to large ranges (such as the layout of the entire working area). The multi - scale sampling module samples the regions related to the initial anchor points at different scales, captures more detailed features, and enables the anchor points to more accurately locate the key features of live working. For example, it may more precisely locate the position where the operator's hand operates the tool at a small scale, and determine the relative position relationship between the operator and the surrounding energized equipment at a large scale, thereby refining the initial anchor points to obtain preliminarily optimized finer - grained anchor points.

[0055] The pyramid spatial attention module further optimizes the finer - grained anchor points. In the live working scenario, certain regions are crucial for accurately recognizing working behaviors, such as the interaction region between the operator and the energized equipment. The pyramid spatial attention module can enhance the extraction and analysis of the features of these regions, further adjust the positions and feature descriptions of the finer - grained anchor points, and make the anchor points more conform to the actual live working behavior features. For example, it can highlight the finer - grained anchor points at the equipment interface where the operator is performing a connection operation, and weaken the influence of the anchor points in some unimportant background regions of the image.

[0056] The task - aware decoupling module performs the final iterative optimization. This module focuses on the specific requirements of the current live working behavior recognition task, and decouples the features closely related to the task from the complex image features. For example, in the task of recognizing whether the live working behavior is compliant, this module will readjust the finer - grained anchor points according to the feature pattern of compliant operations, remove the interference factors unrelated to the task, and strengthen the feature anchor points directly related to the compliance judgment. After this series of iterative optimizations successively performed by the multi - scale sampling module, the pyramid spatial attention module, and the task - aware decoupling module, finally, the optimized finer - grained anchor points are obtained. These optimized finer - grained anchor points can more accurately locate and describe the key features in live working behavior, significantly improving the accuracy and reliability of the live working behavior recognition model.

[0057] Step S250: Topologically align the optimized fine-grained anchors and the encoded live working features to obtain the target live working features.

[0058] In some embodiments, topological alignment refers to establishing a correspondence between the optimized fine-grained anchors and the encoded live working features. On the one hand, in terms of the spatial dimension, it is necessary to ensure that the region located by the optimized fine-grained anchors is accurately associated with the semantic information of the corresponding region in the encoded live working features. For example, in a live working image, if the optimized fine-grained anchor marks the position of a certain equipment interface being operated by a worker, the topological alignment will correspond the semantic description of the operation of this equipment interface (such as the connection state of the interface and the feature encoding of the operation action) in the encoded live working features. On the other hand, at the semantic level, the topological alignment will match and integrate the semantic properties attached to the fine-grained anchors (such as the region where the anchor is located represents the equipment installation operation area) with the relevant semantic information (such as the action process of equipment installation operation and the semantic encoding of tool use, etc.) in the encoded live working features, so that the two echo and cooperate with each other semantically. This topological alignment is not a simple one-to-one correspondence, but comprehensively considers the complex mapping relationship between space and semantics to achieve the deep integration of the two.

[0059] Step S260: Based on the target live working features, perform behavior prediction on the target live working image data to obtain a behavior prediction result.

[0060] The method for predicting live working behaviors in a distribution network based on HOI provided by the present invention first performs multi-scale feature extraction on the acquired target live working image data in the distribution network to obtain multi-scale live working features; performs semantic encoding on the multi-scale live working features to obtain encoded live working features; embeds an initial anchor point in the decoder of the trained live working behavior recognition model, and sequentially iteratively optimizes the fine-grained anchor point through the multi-scale sampling module, the pyramid spatial attention module, and the task-aware decoupling module in the trained live working behavior recognition model to obtain an optimized fine-grained anchor point; performs topological alignment on the optimized fine-grained anchor point and the encoded live working features to obtain target live working features; and performs behavior prediction on the target live working image data based on the target live working features to obtain a behavior prediction result. Thus, the present invention proposes an end-to-end distribution network safety risk recognition model based on Transformer, inserts a lightweight attention module after the multi-scale features output by its backbone network, which can suppress background noise and highlight key regions such as the human body and tools; and improves the encoder and decoder of Transformer, optimizes the computational efficiency of the attention mechanism, and proposes a pyramid spatial attention module, a task-aware decoupling fusion module, and an HOI embedding module to allocate spatial weights to feature maps of different scales, enhance multi-scale spatial relationships, and design independent attention branches for interactive actions and object detection respectively, effectively improving the detection accuracy of predicting live working behaviors in the distribution network and being verified in experiments.

[0061] In some embodiments, the above step S210 further includes the following steps S211 to S214:

[0062] Step S211, acquiring initial live working image data in the distribution network.

[0063] Step S212, removing images with severe occlusion and images that do not contain HOI embodiments in the initial live working image data to obtain the removed live working image data.

[0064] Step S213, annotating the removed live working image data to obtain live working image data annotated with HOI information.

[0065] Step S214, performing data augmentation on the live working image data annotated with HOI information to obtain the target live working image data.

[0066] In some embodiments, the above step S220 is further implemented as follows:

[0067] Based on the multi-scale feature extraction layer in the trained live working behavior recognition model, multi-scale features are extracted from the target live working image data to obtain the multi-scale live working features.

[0068] Among them, the multi-scale feature extraction layer consists of a hierarchical backbone network and a deformable transformer encoder, and the formula is as follows: In the formula, F e ncoder represents the encoder, F f latten represents the flattening operation, φ(·) represents the backbone network, p is the position encoding, s is the spatial shape of the multi-scale feature, r represents the effective ratio, and l represents the level index corresponding to the multi-scale feature. represents a real matrix space with N rows and Cd columns, where N is the number of features and Cd is the dimension of the features.

[0069] In some embodiments, the above step S240 further includes the following steps S241 to S242:

[0070] Step S241, during the decoding process, by establishing a spatial attention region, the decoder performs selective feature aggregation on the target region and maps the optimized semantic instance embedding vector to the HOI embedding.

[0071] Step S242, through the multi-scale sampling module, pyramid spatial attention module, and task-aware decoupling module in the trained live working behavior recognition model for the distribution network, the fine-grained anchor points are iteratively optimized in sequence to obtain the optimized fine-grained anchor points.

[0072] In some embodiments, the HOI-based live working behavior prediction method for the distribution network is applied to the trained live working behavior recognition model, and the trained live working behavior recognition model is trained through the following steps:

[0073] Input the image data of live working on samples into the initial live working behavior recognition model; through the feature extraction layer of the initial live working behavior recognition model, perform multi-scale feature extraction on the image data of live working on samples to obtain multi-scale live working features of samples; through the deformable Transformer encoding layer of the initial live working behavior recognition model, perform semantic encoding on the multi-scale live working features of samples to obtain encoded live working features of samples; embed initial sample anchor points in the decoder of the initial live working behavior recognition model; through the multi-scale sampling module, pyramid spatial attention module and task-aware decoupling module of the initial live working behavior recognition model, iteratively optimize the sample fine-grained anchor points in sequence to obtain optimized sample fine-grained anchor points; perform topological alignment on the optimized sample fine-grained anchor points and the encoded live working features of samples to obtain target sample live working features; through the behavior prediction layer of the initial live working behavior recognition model, based on the target sample live working features, perform behavior prediction on the image data of target sample live working to obtain sample behavior prediction results; input the sample behavior prediction results into a preset loss model to obtain loss results; based on the loss results, correct the model parameters in the initial live working behavior recognition model to obtain a trained live working behavior recognition model.

[0074] Next, an exemplary application of the embodiments of the present invention in a practical application scenario will be described.

[0075] This embodiment proposes a method for predicting live working behavior of a distribution network based on HOI, as Figure 3 shown, including the following steps:

[0076] Step 1, obtain images related to climbing operations in the distribution network as a data set, process the images in the data set and divide the data set into a training set and a test set.

[0077] Further, the processing of the images in the data set specifically includes screening, annotation, and data augmentation.

[0078] Here, screening specifically means removing images in the data set with severe occlusion and not containing HOI embodiments.

[0079] Here, annotation is performed manually. The positions of people and ladders in the images are annotated. Special image annotation tools such as LabelImg software can be used for annotation. Draw bounding boxes on the images and add category labels to identify people and ladders. Then modify according to the data format requirements of the algorithm, including the name of the image file, the identifier of the image, the size of the image, the annotation information of the target objects in the image, and the annotation information of human-object interaction (HOI).

[0080] Here, data augmentation expands the dataset through operations such as rotation, flipping, cropping, scaling, color transformation, adding noise, etc., while updating the bounding boxes to maintain consistency with the targets.

[0081] Step 2: For the end-to-end HOI set prediction model based on Transformer, the training set is input into the model for training to obtain the live working behavior recognition model in the distribution network.

[0082] Furthermore, in the prediction of live working behaviors in the distribution network by this model, first, the input image extracts multi-scale features through a pre-trained backbone network. These features are input into the deformable Transformer encoder for semantic encoding after projection and encoding. The semantic features output by the encoder are fed into the decoder. The decoder generates initial anchor points based on the position embedding, samples from the encoded features using the multi-scale sampling mechanism, and combines the pyramid spatial attention module and the task-aware decoupling fusion module to generate the HOI embedding. Finally, based on the multi-level semantic embedding features output by the encoder and the dynamic anchor boxes constructed by the decoder, the HOI detector infers the human bounding box, object detection box, object category label, and interaction verb semantics, and finally completes the end-to-end parsing of the human-object interaction relationship.

[0083] Furthermore, the multi-scale feature extractor consists of a hierarchical backbone network and a deformable transformer encoder for feature extraction, and the formula is as follows:

[0084]

[0085] where M represents the encoded input features, F e ncoder represents the encoder, F f latten represents the flatten operation, φ(·) represents the backbone network, p is the position encoding, s is the spatial shape of the multi-scale features, r represents the effective ratio, and l represents the level index corresponding to the multi-scale features. represents a real matrix space with N rows and Cd columns, where N is the number of features and Cd is the dimension of the features.

[0086] Here, the hierarchical backbone network is flexible and can be composed of any convolutional neural network and transformer backbone network.

[0087] Furthermore, a lightweight attention module CBAM (Convolutional Block Attention Module) is added to the obtained multi-scale features to perform partition calculation on the obtained multi-scale features, enhance the attention of important features, so as to achieve the effect of suppressing background noise, and the formula is as follows:

[0088]

[0089] Among them, F represents the feature map which contains various feature information of the image, SAM represents the spatial attention module that calculates the attention in the spatial dimension of the input features, CAM(F) is obtained by processing the feature map F to get the activation mapping related to the category, represents the new feature map obtained after the above series of operations, and F' is the output obtained by enhancing and screening the original feature map F.

[0090] Furthermore, during the decoding process, the fine-grained anchor, as a geometric guidance mechanism, by establishing a spatially sensitive attention region, enables the decoder to perform selective feature aggregation on the target region and maps the optimized semantic instance embedding vector to the HOI embedding. This mechanism realizes the noise filtering of feature information through a spatially constrained attention mask, ensuring that the visual representation captured by the content embedding focuses on the discriminative human-object interaction instance features to reduce the interference of non-significant backgrounds. Then, a topological constraint for multi-scale feature alignment is constructed, and the anchor is used to guide the multi-level feature pyramid sampling strategy to optimize the geometric consistency between the query embedding vector and the multi-modal semantic space of the input scene.

[0091] The formulaic definition of the decoding process is as follows:

[0092]

[0093] Among them, H represents the output result of the decoding process, Defattn represents the deformable attention, and Task represents the task-specific module. represents the hierarchical spatial module used to process the spatial structure of the input features and extract multi-scale features in a hierarchical manner, C u represents the content embedding updated by the position embedding. represents the feature sampled at the i-th level.

[0094] Furthermore, compared with the existing method that directly uses the query embedding to generate fine-grained anchors based on the initial anchor without considering the multi-scale features of the input scene and the semantic alignment between the query embedding and the input features, the present invention proposes a novel fine-grained anchor generator, including multi-scale sampling, pyramid spatial attention module, and task-aware decoupled fusion module. This generator makes full use of the initial anchor, multi-scale features, and query embedding to generate appropriate fine-grained anchors for diverse input scenes and aligns the semantic information between different input scenes and query embeddings.

[0095] Furthermore, its multi-scale sampling mechanism, when sampling multi-scale features using the initial anchor points, is mainly used for detecting shallow features of small-sized instances, and the sampling strategy only samples a small range of features around the initial anchor points. For deep features mainly used for detecting large-sized instances, the sampling strategy samples a large range of features around the initial anchor points. Therefore, based on the initial anchor points, the generator uses the sampling strategy to sample multi-scale features as follows:

[0096]

[0097] Among them, F sample represents the feature sampling operation, which is used to extract features of specific regions or scales from the input features. size i (i = 0, 1, 2) represents the sampling size of the i-th level of features. reshape(M) i represents the i-th level of features after reshaping the encoded input feature M. bilinear is bilinear interpolation, which is used to generate a smooth feature map during the upsampling or downsampling process of the feature map. A is the initial anchor point.

[0098] Furthermore, its pyramid spatial attention module uses the spatial attention mechanism (such as channel-separated spatial convolution or non-local operation) to dynamically allocate weights, enhance the response of key regions such as the human body and tools, suppress background interference, and can better utilize the hierarchical spatial information of the sampled features to align the content embedding with the sampled features. The content embedding is first updated through positional embedding and multi-head self-attention mechanism as follows:

[0099] C u = C + F MHA (C + P)W q , (C + P)W k , CW v ;

[0100] Among them, W q , W k and W v respectively represent the query parameter matrix, key parameter matrix, and value parameter matrix in the self-attention mechanism; F MHA is the multi-head attention mechanism; C and P respectively represent the content embedding and positional embedding.

[0101] Then, the updated content embedding is used to merge the sampled features, and the formula is as follows:

[0102]

[0103] Among them, represents the merged feature of the i-th level of sampled features, W OThe parameter matrix representing the multi-head connection, where head1 represents the output feature of the first attention head after being processed by the attention mechanism. represents the output feature of the last attention head, F concat is the concatenation operation. N hd represents the hidden dimension, N H represents the number of attention heads. and represent the query parameter matrix, key parameter matrix, and value parameter matrix of the nth attention head; T represents the transpose operation, which transposes the matrix for matrix multiplication.

[0104] After merging the sampled features of each scale according to the spatial information, the merged features of each scale are first concatenated together as follows:

[0105]

[0106] where X m represents the concatenated multi-scale features merged by the scale-aware merging mechanism. represents the merged feature of the 0th level sampled feature. represents the merged feature of the 2nd level sampled feature, B represents the batch size, that is, the number of samples processed at one time, N q represents the number of branches in the model for different tasks, N L represents the number of multi-scales as follows:

[0107] X u = F concat (head1,..., head h )W O ;

[0108] where X u is the merged multi-scale feature for updating the content embedding, head h represents the output feature of the hth attention head after being processed by the attention mechanism.

[0109] Furthermore, its task-aware decoupled fusion module is used to fuse the merged multi-scale features and content embedding from the task-aware perspective, designs independent attention branches for action recognition and object detection, and avoids feature interference between tasks. Dynamically adjusts the fusion weights of different task branches according to the complexity of the current sample, aligns the content embedding with the merged features, and fuses the content embedding and multi-scale information. The formula is as follows:

[0110]

[0111] Among them, X represents the final feature obtained after being processed by the task-aware decoupling and fusion module, and F stack represents stacking multiple feature vectors together to form a feature matrix.

[0112] After that, the cross-attention mechanism is used to update these features as follows:

[0113] X switch = F concat (head1,..., head h )W O ;

[0114] Among them, Then, the generated information is used to obtain the merged dynamic switch, and the formula is as follows:

[0115]

[0116] Among them, d k represents N hd represents the hidden dimension, N H represents the number of attention heads, Switch γ is the γ-dimensional dynamic switch of the merged features, X switch represents the merged features, which are used to generate the dynamic switch weights. Fnormalize and Fmlp represent the hard sigmoid and the feed-forward network respectively. The feed-forward network consists of two linear layers and one Relu activation layer. The merging mechanism is designed as follows:

[0117]

[0118] Among them, U γ is the γ feature of the content embedding updated by merging multi-scale features; represents the γ feature of the content embedding updated by the position embedding.

[0119] Step 3: Input the test set into the HOI set prediction model for testing.

[0120] Step 4: Compare the HOI set prediction algorithm proposed in this paper with Hog+SVM, Swim transformer, and EfficientnetV2.

[0121] Furthermore, according to the evaluation criteria, the role average precision is used to evaluate the predicted HOI instances. If the intersection over union (IOU) of the detected bounding box and the true annotation bounding box of the same category is greater than 0.5, the object detection is considered a true positive example.

[0122] Furthermore, on the dataset of ladder climbing, this algorithm was compared with the Hog+SVM, Swim transformer, and EfficientnetV2 algorithms. The mAP was significantly improved by 5.85%, and the model can accurately identify HOI instances from a noisy background and can withstand the impact of harsh environments and the recognition of small targets.

[0123] Figure 4 It is a schematic diagram of the composition structure of the device for predicting live working behavior of a distribution network based on HOI provided by an embodiment of the present invention. As Figure 4 shown, the device 400 for predicting live working behavior of a distribution network based on HOI includes: an acquisition module 401 for acquiring target live working image data in the distribution network; an extraction module 402 for performing multi-scale feature extraction on the target live working image data to obtain multi-scale live working features; an encoding module 403 for performing semantic encoding on the multi-scale live working features based on the deformable Transformer encoder in the trained live working behavior recognition model to obtain encoded live working features; an optimization module 404 for embedding an initial anchor point in the decoder of the trained live working behavior recognition model and iteratively optimizing the fine-grained anchor point through the multi-scale sampling module, the pyramid spatial attention module, and the task-aware decoupling module in the trained live working behavior recognition model to obtain an optimized fine-grained anchor point; an alignment module 405 for performing topological alignment on the optimized fine-grained anchor point and the encoded live working features to obtain target live working features; and a prediction module 406 for performing behavior prediction on the target live working image data based on the target live working features to obtain a behavior prediction result.

[0124] It should be noted that the description of the device in the embodiment of the present invention is similar to the description of the above method embodiment and has the same beneficial effects as the method embodiment, so it will not be repeated. For the technical details not disclosed in this device embodiment, please refer to the description of the method embodiment of the present invention for understanding.

[0125] Based on the foregoing embodiments, an embodiment of the present invention also provides an electronic device. Figure 5 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. As Figure 5 shown, the hardware entity of the electronic device 500 includes: a memory 501 and a processor 502. The memory 501 stores a computer program that can run on the processor 502, and when the processor 502 executes the program, it implements the steps in the prediction of live working behavior of a distribution network based on HOI in the above embodiments.

[0126] The memory 501 is configured to store instructions and applications executable by the processor 502, and can also cache data to be processed or already processed by the processor 502 and each module in the electronic device 500 (e.g., image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0127] Based on the foregoing embodiments, an embodiment of the present invention further provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor of an electronic device, it can implement the method for predicting live working behaviors of a distribution network based on HOI provided in any of the foregoing embodiments.

[0128] The descriptions of the foregoing embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated herein.

[0129] The methods disclosed in the method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0130] The features disclosed in the product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0131] The features disclosed in the method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0132] It should be noted that the above computer-readable storage medium can be a ferroelectric memory (FRAM, Ferromagnetic Random Access Memory), read-only memory (ROM, Read Only Memory), programmable read-only memory (PROM, Programmable Read Only Memory), erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory), electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read Only Memory), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM, Compact Disk-Read Only Memory), etc.; it can also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0133] It should be noted that, in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method or apparatus including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method or apparatus. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or apparatus including such element. In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or integrated into another system, or some features can be ignored, or not executed.

[0134] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus necessary general hardware nodes. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0136] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0137] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0139] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for predicting live working behaviors in a distribution network based on HOI, characterized in that, The method includes: Obtain target live working image data in the distribution network; Perform multi-scale feature extraction on the target live working image data to obtain multi-scale live working features; Based on the deformable Transformer encoder in the trained live working behavior recognition model, perform semantic encoding on the multi-scale live working features to obtain encoded live working features; Embed an initial anchor point in the decoder of the trained live working behavior recognition model, and sequentially iterate and optimize the fine-grained anchor points through the multi-scale sampling module, pyramid spatial attention module, and task-aware decoupling module in the trained live working behavior recognition model to obtain optimized fine-grained anchor points; Perform topological alignment on the optimized fine-grained anchor points and the encoded live working features to obtain target live working features; Based on the target live working features, perform behavior prediction on the target live working image data to obtain a behavior prediction result.

2. The method according to claim 1, wherein The obtaining of the target live working image data in the distribution network includes: Obtain initial live working image data in the distribution network; Remove the images with serious occlusion and without HOI examples in the initial live working image data to obtain the removed live working image data; Annotate the removed live working image data to obtain live working image data with annotated HOI information; Perform data augmentation on the live working image data with annotated HOI information to obtain the target live working image data.

3. The method according to claim 1, characterized in that, The performing of multi-scale feature extraction on the target live working image data to obtain multi-scale live working features includes: Based on the multi-scale feature extraction layer in the trained live working behavior recognition model, perform multi-scale feature extraction on the target live working image data to obtain the multi-scale live working features; Among them, the multi-scale feature extraction layer consists of a hierarchical backbone network and a deformable transformer encoder, and the formula is as follows: In the formula, F e ncoder represents the encoder, F f latten represents the flattening operation, φ() represents the backbone network, p is the position encoding, s is the spatial shape of the multi-scale feature, r represents the effective ratio, and l represents the level index corresponding to the multi-scale feature. represents a space of real number matrices with N rows and Cd columns, where N is the number of features and Cd is the dimension of the features.

4. The method according to claim 1, wherein The embedding of the initial anchor point in the decoder of the trained live working behavior recognition model and the sequential iteration and optimization of the fine-grained anchor points through the multi-scale sampling module, pyramid spatial attention module, and task-aware decoupling module in the trained distribution network live working left-right behavior recognition model to obtain optimized fine-grained anchor points includes: During the decoding process, by establishing a spatial attention region, enabling the decoder to perform selective feature aggregation on the target region, and mapping the optimized semantic instance embedding vector to HOI embedding; Sequentially iterate and optimize the fine-grained anchor points through the multi-scale sampling module, pyramid spatial attention module, and task-aware decoupling module in the trained distribution network live working left-right behavior recognition model to obtain optimized fine-grained anchor points.

5. The method according to claim 1, wherein The HOI-based distribution network live working behavior prediction method is applied to a trained live working behavior recognition model, and the trained live working behavior recognition model is trained through the following steps: Input sample live working image data into the initial live working behavior recognition model; Through the feature extraction layer of the initial live working behavior recognition model, perform multi-scale feature extraction on the sample live working image data to obtain sample multi-scale live working features; Semantically encode the sample multi-scale live working features through the deformable Transformer encoding layer of the initial live working behavior recognition model to obtain the encoded sample live working features; Embed the initial sample anchors in the decoder of the initial live working behavior recognition model; Iteratively optimize the sample fine-grained anchors through the multi-scale sampling module, pyramid spatial attention module and task-aware decoupling module of the initial live working behavior recognition model to obtain the optimized sample fine-grained anchors; Topologically align the optimized sample fine-grained anchors and the encoded sample live working features to obtain the target sample live working features; Based on the target sample live working features, perform behavior prediction on the target sample live working image data through the behavior prediction layer of the initial live working behavior recognition model to obtain the sample behavior prediction results; Input the sample behavior prediction results into a preset loss model to obtain the loss results; Based on the loss results, correct the model parameters in the initial live working behavior recognition model to obtain the trained live working behavior recognition model.

6. A live working behavior prediction device for a distribution network based on HOI, characterized in that, The device includes: An acquisition module for acquiring target live working image data in the distribution network; An extraction module for performing multi-scale feature extraction on the target live working image data to obtain multi-scale live working features; An encoding module for semantically encoding the multi-scale live working features based on the deformable Transformer encoder in the trained live working behavior recognition model to obtain the encoded live working features; An optimization module for embedding initial anchors in the decoder of the trained live working behavior recognition model and iteratively optimizing the fine-grained anchors through the multi-scale sampling module, pyramid spatial attention module and task-aware decoupling module of the trained live working behavior recognition model to obtain the optimized fine-grained anchors; An alignment module for topologically aligning the optimized fine-grained anchors and the encoded live working features to obtain the target live working features; A prediction module for performing behavior prediction on the target live working image data based on the target live working features to obtain the behavior prediction results.

7. An electronic device, characterized in that, It includes: A memory for storing executable instructions; A processor for implementing the method for predicting live working behavior in a distribution network based on HOI according to any one of claims 1 to 5 when executing the executable instructions stored in the memory.

8. A computer-readable storage medium, characterized in that, Stored with executable instructions for causing the processor to implement the method for predicting live working behavior in a distribution network based on HOI according to any one of claims 1 to 5 when executing the executable instructions.