Power operation violation identification method and system based on key point constraint
By transforming safety operation regulations into constraint relationships between key points, combining graphic inference models and image analysis technology, the problem of traditional artificial intelligence being difficult to identify complex violations is solved, and accurate identification and automated monitoring of violations in power operations is achieved, and safety and efficiency are improved.
Patent Information
- Application Number
- CN202510256050.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional artificial intelligence image recognition technology is difficult to identify violations in power operations in complex scenarios, especially the inability to couple visual features with safety regulations and text information for analysis. It has a narrow scope of application and poor effect.
The power operation violation identification method based on key point constraints is adopted. By transforming safety operation regulations into constraint relationships between key points, the graphic and text reasoning model can be analyzed in combination with operation safety procedures and on-site monitoring images to achieve accurate analysis and judgment of violations.
It improves the accuracy of identification of violations, realizes automatic identification, reduces the work burden of power safety supervision personnel, improves safety supervision efficiency, and can monitor and warn in real time, effectively reducing the occurrence of high-risk violations.
Smart Images

Figure CN120182913A_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a method and system for identifying violations of electric power operations based on key point constraints, and relates to the field of safety management and control of power grid projects. Background Art
[0002] The on-site operation points of power grid projects are numerous and wide-ranging, involving daily operation and maintenance, equipment maintenance, infrastructure construction and other aspects. In addition, the operating environment is complex, and the overall management and control is difficult. In order to further improve the level of safe production, the company proposed the "four controls" operation risk management strategy. "Four controls" means "control the plan, control the team, control the personnel, and control the site". Among them, personnel are the key to the implementation of operation risk management measures, and the site is the core of risk management and safety measures. The constraints and supervision of these two elements can effectively reduce the operation risks. However, at present, the power grid operators and operation sites are widely distributed and numerous, making it difficult to cover and monitor them with manpower. Therefore, the implementation of "four controls" cannot be separated from the support of digital technology.
[0003] In order to promote the application of new technologies of "safety supervision + artificial intelligence" and realize the digital transformation of the company's safety supervision, the Safety Supervision Department compiled the "Safety Supervision Professional Artificial Intelligence Application Work Plan" and organized related industries to carry out research and implementation of intelligent terminal technologies for digital safety management and control of on-site operations. At present, the company has preliminarily completed the digital and intelligent evolution of identifying violations at the work site. However, in the actual application of related technologies, there are still problems that need to be solved. Among them, the more typical ones are: traditional artificial intelligence image recognition technology can usually only identify violations with obvious visual features in simple scenarios, such as not wearing a safety helmet, not dressing as required, etc., because the model can only complete the task based on the visual features of the image, and it is impossible to couple it with text information such as safety regulations and typical violation libraries to complete the analysis of violations. The scope of application is narrow and the application effect is poor. Summary of the invention
[0004] In response to the above problems, the present invention proposes a method and system for identifying violations in power operations based on key point constraints. The method converts safety operating regulations into constraint relationships between key points, so that the graphic reasoning model can intuitively combine the requirements of operating safety regulations with on-site monitoring images for analysis, thereby achieving accurate judgment of violations.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is: In a first aspect, the present invention provides a method for identifying violations of power operations based on key point constraints, comprising: Obtain on-site operation videos, images, and operation type text; Based on the on-site operation videos, images, and operation type texts, use a graphic-text reasoning model to infer the tools to be detected and the corresponding part names; send the part names into an image semantic extraction model and an image segmentation model respectively to extract the semantic block features and masks of each part of the tool; Upsample the semantic block features to the original image size through bilinear interpolation; cluster the semantic block features corresponding to each mask based on the original image size, and use the cluster centers as the key point candidates for each part of the tool; project the key point candidates of each part of the tool into the world coordinate system; Extract the key points of each part of the human body according to the on-site operation videos, images, and operation type texts; project the human key point candidates into the world coordinate system; Based on the projections of the key point candidates of each part of the tool and the human key point candidates in the world coordinate system, and according to the point constraint knowledge corresponding to the detailed operation types, determine whether there are any operation violation behaviors.
[0006] As a further improvement of the present invention, the projection of the key point candidates of each part of the tool into the world coordinate system includes: Project the key point candidates of each part of the tool into the 3D world coordinate system through the depth map generated by DINOv2; The projection of the human key point candidates into the world coordinate system includes: Project the human key point candidates into the 3D world coordinate system through the depth map generated by DINOv2.
[0007] As a further improvement of the present invention, after the projection of the key point candidates of each part of the tool into the world coordinate system, it further includes: Filter out the key points with a distance less than the threshold in the world coordinate system.
[0008] As a further improvement of the present invention, before determining whether there are any operation violation behaviors according to the point constraint knowledge corresponding to the detailed operation types, it further includes: When there are multiple operators, according to the principle of proximity, group the human key points and the tool key points, and then perform grouped processing according to the grouping situation.
[0009] As a further improvement of the present invention, the extraction of the key points of each part of the human body according to the on-site operation videos, images, and operation type texts is to use a human pose recognition model to extract the key points of each part of the human body; The human pose recognition model is used to extract the positions of the key parts of the body and can represent the key points of human behavior.
[0010] As a further improvement of the present invention, the determination of whether there is a violation in the operation is based on the graphic reasoning model to determine whether there is a violation in the operation according to the point constraint knowledge corresponding to the detailed operation types.
[0011] As a further improvement of the present invention, the graphic reasoning model is used to determine the detailed operation types according to the input on-site surveillance video and operation type text; to derive the relevant tools and their component names according to the detailed operation types; and to derive whether there is a violation according to the spatial position relationship between the key points of the tools and the key points of the human body. The graphic reasoning model adopts a vision-language large model, and obtains power operation knowledge and tool knowledge through fine-tuning or knowledge retrieval enhancement, and learns key point constraint knowledge to obtain the position relationship between key points and violation behaviors.
[0012] As a further improvement of the present invention, the image semantic extraction model is used to extract the semantic block features of the tools and divide each component of the tool according to its function; the image segmentation model is used to generate masks for each component of the tool to support the subsequent extraction of key points.
[0013] In a second aspect, the present invention provides a power operation violation recognition system based on key point constraints, including: An acquisition module, configured to acquire on-site operation videos, images, and operation type texts. An inference and extraction module, configured to infer the tools to be detected and their corresponding component names by using a graphic reasoning model according to the on-site operation videos, images, and operation type texts; and send the component names to the image semantic extraction model and the image segmentation model respectively to extract the semantic block features and masks of each component of the tool. A clustering and projection module, configured to upsample the semantic block features to the original image size through bilinear interpolation; cluster the semantic block features corresponding to each mask based on the original image size, and use the cluster centers as the key point candidates for each component of the tool; project the key point candidates for each component of the tool into the world coordinate system. An extraction and projection module, configured to extract the key points of each part of the human body according to the on-site operation videos, images, and operation type texts; project the human key point candidates into the world coordinate system. A behavior determination module, configured to determine whether there is a violation in the operation based on the projections of the key point candidates of each component of the tool and the human key point candidates in the world coordinate system according to the point constraint knowledge corresponding to the detailed operation types.
[0014] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for identifying power operation violations based on key point constraints is implemented.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method for identifying power operation violations based on key point constraints.
[0016] In a fifth aspect, the present invention provides a computer program product including computer instructions that direct a computer to execute the method for identifying power operation violations based on key point constraints.
[0017] The beneficial effects of the present invention compared with the prior art are as follows: Traditional violation behavior analysis techniques based on object recognition are often limited by the performance of the object detection model and the diversity of training data. However, this method can more flexibly handle complex and changeable violation behaviors by transforming them into the constraint relationships between points and positions. Through the inference of the graphic reasoning model and deep learning techniques such as image semantic extraction and segmentation, the key points of tools and human bodies can be extracted more accurately, thereby improving the recognition accuracy of violation behaviors. This method realizes the automatic recognition of violation behaviors, which can greatly reduce the workload of power safety supervision personnel and improve the supervision efficiency. At the same time, the automated recognition system can also achieve real-time monitoring and early warning, and promptly discover and correct violation behaviors. It mainly aims at violation behaviors such as not wearing safety belts correctly during high-altitude operations, which have high risks and hazards in power operations. Through targeted recognition methods, the occurrence of such violation behaviors can be effectively reduced, ensuring the safety of operators. The method for identifying power operation violations based on key point constraints of the present invention has clear principles and obvious advantages, and can effectively improve the safety and efficiency of power operations. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings below only conveniently and clearly show some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of a method for identifying power operation violations based on key point constraints of the present invention; Figure 2A flowchart of power operation violation recognition based on key point constraints given by an embodiment of the present invention; Figure 3 A power operation violation recognition device based on key point constraints provided by the present invention; Figure 4 A schematic diagram of an electronic device provided by the present invention. Detailed implementation manners
[0020] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0021] In the description of the present invention, unless otherwise clearly defined, terms such as "set", "installed", and "connected" should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above terms in the present invention in combination with the specific content of the technical solution.
[0022] As Figure 1 shown, the first object of the present invention is to provide a power operation violation recognition method based on key point constraints, including: S1. Obtain on-site operation videos, images, and operation type texts; Further, data acquisition: First, obtain necessary information through on-site operation videos, images, and operation type texts. These information are the basis for subsequent reasoning and recognition.
[0023] S2. According to the on-site operation videos, images, and operation type texts, use a graphic-text reasoning model to infer the tools to be detected and the corresponding part names; send the part names into an image semantic extraction model and an image segmentation model respectively to extract the semantic block features and masks of each part of the tool; Further, graphic-text reasoning model reasoning: Use the graphic-text reasoning model to infer the tools to be detected and the corresponding part names according to the on-site operation videos, images, and operation type texts. This step realizes the fusion of text content and input data, improving the accuracy of reasoning.
[0024] Further, image semantic extraction and segmentation: Send the part names into an image semantic extraction model and an image segmentation model respectively to extract the semantic block features and masks of each part of the tool. The semantic block features reflect the image content of the part, while the mask is used to locate the position of the part in the image.
[0025] S3, Upsample the semantic block features to the original image size through bilinear interpolation; cluster the semantic block features corresponding to each mask based on the original image size, and use the cluster centers as candidate key points for each component of the tool; project the candidate key points of each component of the tool into the world coordinate system; Further, feature upsampling and clustering: Upsample the semantic block features to the original image size through bilinear interpolation, and then cluster the semantic block features corresponding to each mask. The cluster centers are used as candidate key points for each component of the tool. These candidate key points represent the key positions of the components in the image.
[0026] S4, Extract the key points of each part of the human body according to the on-site operation video, images and operation type text; project the candidate human body key points into the world coordinate system; Further, world coordinate system projection: Project the candidate key points of each component of the tool and the candidate human body key points into the world coordinate system. This step realizes the conversion from the image space to the actual space, facilitating subsequent point constraint judgment.
[0027] S5, Based on the projections of the candidate key points of each component of the tool and the candidate human body key points in the world coordinate system, and according to the point constraint knowledge corresponding to the detailed operation type, determine whether there is any operation violation behavior.
[0028] Further, violation behavior determination: According to the point constraint knowledge corresponding to the detailed operation type, determine whether there is any operation violation behavior. The point constraint knowledge is formulated based on safety operation requirements and regulations, and is used to judge whether the positional relationship between key points complies with the regulations.
[0029] The present invention proposes a method for identifying power operation violations based on key point constraints. By converting safety operation requirements and regulations into visible point positions and the constraint relationships between point positions, this method effectively realizes the integration of text content and input data, maximizing the inference performance of the graphic-text reasoning model. This method can break through the limitations of the violation behavior analysis technology based on object recognition, effectively identify complex violation behaviors, and contribute to the promotion of the construction of power safety supervision automation. It mainly aims at violation behaviors such as not wearing a safety belt correctly during high-altitude operations.
[0030] The following further elaborates on the present invention with specific embodiments: As Figure 2 shown, the method for identifying power operation violation behaviors proposed by the present invention includes the following steps: Step 1, Let the graphic-text reasoning model infer the names of the tools and their components to be detected according to the input on-site operation video, images and operation type text.
[0031] In this step, first, videos, images of on-site operations, and text information describing the types of operations are received. These information are input into a pre-trained image-text reasoning model, which can understand and parse the input multi-modal data. Based on the operation type and image content, the model infers the names of various tools and their specific components that may need to be detected and recognized in the current operation scenario. The output of this step is a list containing the names of the components to be detected.
[0032] Step 2: Send the component names into the image semantic extraction model and the image segmentation model respectively to extract the patch-wise semantic block features and masks of each component.
[0033] Next, for each inferred component name, the system sends it into two models respectively: the image semantic extraction model and the image segmentation model. The image semantic extraction model is responsible for extracting the patch-wise (i.e., local area) semantic block features of the component, which can reflect key information such as the texture and shape of the component. The image segmentation model generates a mask corresponding to the component, that is, a binary image, where the component area is white (or 1) and the rest of the area is black (or 0). These two outputs provide key information for the subsequent steps.
[0034] Step 3: Upsample the semantic block features to the size of the original image by bilinear interpolation.
[0035] Furthermore, to align the semantic block features with the original image, the system uses bilinear interpolation to upsample the semantic block features to the size of the original image. This step ensures that the spatial position of the features in the subsequent processing is consistent with the original image.
[0036] Step 4: Perform k-means clustering (k = 5, using cosine similarity metric) on the patch-wise semantic block features corresponding to each mask, and the cluster centers are used as key point candidates.
[0037] Furthermore, for the semantic block features corresponding to each mask, the system uses the k-means clustering algorithm (such as setting k = 5 and using cosine similarity as the metric standard) for clustering. The purpose of clustering is to extract representative key point candidates from the features. Each cluster center represents a local area with similar features, and these centers are regarded as potential key points.
[0038] Step 5: Project the tool key point candidates into the 3D world coordinate system through the depth map generated by DINOv2.
[0039] Further, to more accurately understand the positional relationship of the tools in the three-dimensional space, the system projects the key-point candidates obtained in Step 4 into the 3D world coordinate system using the depth map generated by DINOv2 (a deep neural network model). This step provides the three-dimensional position information of the key points and lays the foundation for subsequent spatial relationship analysis.
[0040] Step 6: Filter out key points that are too close to avoid redundancy (for example, for multiple key points with a distance less than 10 cm, only one can be retained).
[0041] Further, in the three-dimensional space, there may be key points that are very close to each other, and these key points may generate redundant information during recognition and analysis. Therefore, the system sets a distance threshold (for example, 10 cm). For key points with a distance less than this threshold, only one of them is retained as a representative to reduce the complexity of subsequent processing.
[0042] Step 7: Use a human pose recognition model (such as YOLO11) to extract the key points of each part of the human body.
[0043] Further, to analyze the posture and behavior of the operator, the system uses a human pose recognition model (such as a variant of YOLOv11 or a similar model) to extract the key points of each part of the human body (such as the head, shoulders, elbows, hands, etc.) from the image. These key points reflect the posture information of the operator.
[0044] Step 8: Project the human key-point candidates into the 3D world coordinate system through the depth map generated by DINOv2.
[0045] Further, similar to Step 5, the system projects the human key points into the 3D world coordinate system using the depth map generated by DINOv2. This step provides the position information of the operator in the three-dimensional space, facilitating the subsequent spatial relationship analysis with the key points of the tools.
[0046] Step 9: Let the text-image reasoning model determine whether there are any job violation behaviors according to the point constraint knowledge corresponding to the job sub-types.
[0047] Finally, based on the point constraint knowledge corresponding to the job sub-types (these constraint knowledge may be obtained based on rules, expert systems, or machine learning models), the system uses the text-image reasoning model to comprehensively analyze the key points obtained in Step 5 and Step 8. The model determines whether there are any job violation behaviors according to information such as the spatial relationship between the key points, the posture of the operator, and the usage method of the tools. If there is a violation behavior, the system will give corresponding prompts or alarm information.
[0048] To implement the above solution, the method proposed by the present invention involves four types of models, namely, a human pose recognition model, a graphic-text reasoning model, an image semantic extraction model, and an image segmentation model. The following provides a detailed description of each model.
[0049] Human pose recognition model: It is used to extract the positions of key body parts, that is, to provide key points that can represent human behaviors.
[0050] Preferably, the YOLO11 open-source model can be adopted to complete the recognition of key body parts without fine-tuning. By default, the YOLO11 model can extract 17 key points from each person in the picture. The symmetrically distributed key points are: 0 - nose, 1 - left eye, 2 - right eye, 3 - left ear, 4 - right ear, 5 - left shoulder, 6 - right shoulder, 7 - left elbow, 8 - right elbow, 9 - left wrist, 10 - right wrist, 11 - left hip, 12 - right hip, 13 - left knee, 14 - right knee, 15 - left ankle, 16 - right ankle.
[0051] As an alternative solution, the graphic-text reasoning model mainly has three functions: 1) It is used to judge the detailed types of operations based on the input on-site surveillance video and operation type text, such as distinguishing pole climbing, tower climbing, and platform operations.
[0052] 2) Derive the relevant tools and their component names according to the detailed operation types. For example, pole climbing corresponds to a waist belt, safety rope, safety hook, etc.; tower climbing or platform operations correspond to a back rope, speed differential self-locking device, etc.
[0053] 3) Deduce whether there are any violation behaviors based on the spatial position relationship between the key points of the tools and the key points of the human body.
[0054] Preferably, the Qwen2-VL-7B open-source large model can be adopted. The model needs to obtain power operation knowledge and tool knowledge through fine-tuning or knowledge retrieval enhancement to ensure the realization of 1) and 2). At the same time, it needs to learn key point constraint knowledge to clarify the position relationship between key points and the corresponding method of violation behaviors to ensure the realization of 3).
[0055] Regarding the violation behavior of not wearing a safety belt correctly during high-altitude operations, the present invention provides the following key point constraint knowledge in Table 1.
[0056] Table 1
[0057] Furthermore, the key point constraint knowledge can be expanded and updated according to the types of violations to be recognized. The key lies in using the position relationship between the tool components and the human body to represent the operation requirements.
[0058] As an optional solution, image semantic extraction model: used to extract patch-wise (semantic block) features of tools, which can be used as a basis to divide the tool parts according to their functions and support the subsequent extraction of key points. Preferably, the DINOv2 open source model can be used, which can be fine-tuned as needed to provide higher tool part recognition accuracy.
[0059] As an optional solution, the image segmentation model is used to generate masks for each component of the tool to support the subsequent extraction of key points. Preferably, the Segment Anything Model (SAM) series of models can be used, which can be fine-tuned as needed to provide higher tool component segmentation accuracy.
[0060] It should be noted that when there are multiple workers in the surveillance data, it is necessary to group the key points of the human body and the key points of the tools according to the principle of proximity before step 9, and then let the graphic reasoning model process the groups separately.
[0061] As an example, the method of the present invention can be used to process images, video data collected by the surveillance ball or image data stored on the risk control platform to identify illegal behaviors in high-altitude operations. It is only necessary to sort out the relationship between the operator, the operating tool and the operating environment according to the safety regulations and express it as the height and left-right relationship between the key points.
[0062] For example, the method of the present invention determines whether there is any illegal behavior of adjusting and removing the pull wire by identifying the change in the relationship between the free end of the pull wire and the key points of the person's body; for another example, it determines whether there is any crossing behavior by identifying the height relationship between the person's left and right knees, left and right ankles and the safety fence.
[0063] Of course, the method of the present invention can be used to judge various types of work violations, including but not limited to: 1) Failure to wear a safety belt correctly during high-altitude work. 2) Adjusting or removing the guy wire when someone is working on the tower. 3) Workers wearing or crossing safety fences or safety cordons without authorization.
[0064] like Figure 3 As shown, the third object of the present invention is to provide a power operation violation identification system based on key point constraints, comprising: Acquisition module 100, used to acquire on-site operation videos, images and operation type texts; The inference and extraction module 200 is used to infer the tools and the corresponding parts names that need to be inspected based on the on-site operation video, image and operation type text by using the image-text inference model; the part names are respectively sent to the image semantic extraction model and the image segmentation model to extract the semantic block features and masks of each part of the tool; The clustering projection module 300 is used to upsample the semantic block features to the original image size through bilinear interpolation; cluster the semantic block features corresponding to each mask based on the original image size, and use the cluster centers as the key point candidates for each component of the tool; project the key point candidates for each component of the tool into the world coordinate system; The extraction projection module 400 is used to extract the key points of each part of the human body according to the on-site operation video, images, and operation type text; project the human key point candidates into the world coordinate system; The behavior discrimination module 500 is used to determine whether there is an operation violation behavior based on the projections of the key point candidates for each component of the tool and the human key point candidates in the world coordinate system, according to the point constraint knowledge corresponding to the detailed operation types.
[0065] The power operation violation recognition system based on key point constraints of the present invention is based on the above-mentioned power operation violation recognition method based on key point constraints.
[0066] As Figure 4 shown, the third object of the embodiment of the present invention is to provide an electronic device, including a memory 701, a processor 702, and a computer program stored in the memory 701 and executable on the processor. When the processor executes the computer program, the above-mentioned power operation violation recognition method based on key point constraints is implemented. It also includes a communication interface 703 and a bus 704.
[0067] The fourth object of the embodiment of the present invention is to provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned power operation violation recognition method based on key point constraints is implemented.
[0068] The fifth object of the embodiment of the present invention is to provide a computer program product, which includes computer instructions. The computer instructions instruct the computer to execute the above-mentioned power operation violation recognition method based on key point constraints.
[0069] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or Figure 1 boxes or multiple boxes.
[0070] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or steps for implementing the functions specified in multiple blocks or blocks.
[0071] The present invention may be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, readable storage media, optical storage, etc.) containing computer-usable program code.
[0072] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks or blocks.
[0073] Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for identifying violations of power operations based on key point constraints, characterized in that: include: Obtain on-site operation videos, images, and operation type text; Based on the on-site operation video, images and operation type text, the graph-text reasoning model is used to infer the tools and corresponding parts names that need to be inspected; the part names are sent to the image semantic extraction model and image segmentation model respectively to extract the semantic block features and masks of each part of the tool; The semantic block features are upsampled to the original image size through bilinear interpolation; the semantic block features corresponding to each mask are clustered based on the original image size, and the cluster centers are used as key point candidates for each component of the tool; the key point candidates for each component of the tool are projected into the world coordinate system; Extract key points of various parts of the human body based on on-site operation videos, images and operation type texts; Project the human key point candidates into the world coordinate system; Based on the projection of the key point candidates of each part of the tool and the key point candidates of the human body into the world coordinate system, and according to the point constraint knowledge corresponding to the subdivision type of the operation, it is determined whether there is any violation of the operation.
2. The method for identifying violations of power operations based on key point constraints according to claim 1 is characterized in that: The projecting of the key point candidates of each component of the tool into the world coordinate system includes: Project the key point candidates of each part of the tool into the 3D world coordinate system through the depth map generated by DINOv2; The projecting of the human body key point candidates into the world coordinate system includes: The depth map generated by DINOv2 is used to project the human key point candidates into the 3D world coordinate system.
3. The method for identifying violations of power operations based on key point constraints according to claim 1 is characterized in that: After projecting the key point candidates of each component of the tool to the world coordinate system, the method further includes: Filter out keypoints whose distance in world coordinates is less than a threshold.
4. The method for identifying violations of power operations based on key point constraints according to claim 1 is characterized in that: Before judging whether there is any violation of operation rules based on the point constraint knowledge corresponding to the operation subdivision type, the method further includes: When there are multiple operators, the key points of the human body and the key points of tools are grouped according to the principle of proximity, and then the groups are processed according to the grouping situation.
5. The method for identifying violations of power operations based on key point constraints according to claim 1 is characterized in that: The method of extracting key points of various parts of the human body based on the on-site operation video, image and operation type text is to extract key points of various parts of the human body using a human posture recognition model; The human body posture recognition model is used to extract the positions of key parts of the body and can characterize the key points of human behavior.
6. The method for identifying violations of power operations based on key point constraints according to claim 1 is characterized in that: The determination of whether there is an operation violation is based on a graph-text reasoning model according to point constraint knowledge corresponding to the operation subdivision category to determine whether there is an operation violation.
7. The method for identifying violations of power operations based on key point constraints according to claim 1 is characterized in that: The graphic reasoning model is used to determine the subdivision of the operation based on the input on-site monitoring video and operation type text; to deduce the names of related tools and their components based on the subdivision of the operation; and to deduce whether there is a violation of regulations based on the spatial position relationship between the key points of the tools and the key points of the human body; The graphic reasoning model adopts a large visual language model, acquires power operation knowledge and tool knowledge through fine-tuning or knowledge retrieval enhancement, and learns key point constraint knowledge to obtain the positional relationship and violation behavior between key points.
8. The method for identifying violations of power operations based on key point constraints according to claim 1 is characterized in that: The image semantic extraction model is used to extract the semantic block features of the tool and divide the various parts of the tool according to their functions; the image segmentation model is used to generate masks for the various parts of the tool to support the subsequent extraction of key points.
9. A power operation violation identification system based on key point constraints, characterized in that: include: The acquisition module is used to obtain on-site operation videos, images and operation type texts; The inference and extraction module is used to infer the tools and corresponding parts names that need to be inspected based on the on-site operation video, images and operation type text using the image-text inference model; the part names are respectively sent to the image semantic extraction model and the image segmentation model to extract the semantic block features and masks of each part of the tool; A clustering projection module is used to upsample the semantic block features to obtain the original image size through bilinear interpolation; cluster the semantic block features corresponding to each mask based on the original image size, and use the cluster center as the key point candidate of each component of the tool; and project the key point candidate of each component of the tool into the world coordinate system; Extraction and projection module, used to extract key points of various parts of the human body based on on-site operation videos, images and operation type texts; Project the human key point candidates into the world coordinate system; The behavior discrimination module is used to judge whether there is any violation of the operation based on the projection of the key point candidates of each part of the tool and the key point candidates of the human body into the world coordinate system and the point constraint knowledge corresponding to the operation subdivision type.
10. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for identifying violations of electric power operations based on key point constraints as described in any one of claims 1 to 8 when executing the computer program.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for identifying violations of electric power operations based on key point constraints as described in any one of claims 1 to 8 is implemented.
12. A computer program product, comprising computer instructions, characterized in that: The computer instructions instruct the computer to execute the method for identifying violations of power operations based on key point constraints as described in any one of claims 1-8.