Automatic detection method for power safety violation behaviors
Through knowledge graph-driven multimodal data construction and adversarial sample generation technology, the problems of insufficient scene coverage and weak cross-modal semantic correlation in power safety detection are solved, and high-precision detection of composite violations is achieved.
Patent Information
- Application Number
- CN202510577571.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing power safety detection methods have problems such as insufficient scene coverage and weak cross-modal semantic correlation in complex power operation environments, resulting in insufficient identification accuracy of compound violations.
Through knowledge graph-driven multimodal data construction and dynamic semantic recombination technology, an image-text alignment data set covering the entire scene of power operations is generated, and a dual-modal adversarial sample generation technology is introduced to enhance the robustness of the model.
It significantly improves the scene coverage ability and cross-modal semantic understanding accuracy of complex violations, enhances the model's resistance to visual noise and semantic misleading, and improves the accuracy of detection of composite violations.
Smart Images

Figure CN120088864A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and deep learning multimodality, etc., and in particular to an automatic detection method for power safety violation behaviors. Background Art
[0002] With the continuous improvement of the digital upgrade of the power industry and the demand for intelligent operation and maintenance, the safety monitoring of power operations has gradually shifted from traditional manual inspections to intelligent technology-driven. Although object detection technology based on computer vision has achieved certain application results in industrial scenarios, its practice in the field of power safety still faces significant challenges: existing detection methods mostly rely on general-scenario training models and are difficult to adapt to the unique complex environment and refined specification requirements of power operations. For example, in scenarios such as the detection of the wearing of high-altitude operation protection equipment and the identification of the compliance of high-voltage equipment operations, traditional algorithms have insufficient recognition accuracy for multi-view occlusion, dynamic light changes, and subtle violation behaviors, and lack in-depth integration of professional field knowledge such as the "Power Safety Work Regulations". At the same time, there are structural defects at the data level. Existing data sets generally lack multi-modal annotation information covering the entire power scenario, and the semantic association between images and text descriptions is weak, making it difficult to support the cross-modal association learning ability required by large models, resulting in limited understanding and reasoning ability of the model for complex violation scenarios (such as "not wearing a seat belt and not setting up an insulating partition") during actual deployment. This double lack of data and knowledge severely restricts the reliability and interpretability of intelligent detection systems in complex power operation environments. Summary of the Invention
[0003] Aiming at the defects and deficiencies of the existing technologies, the present invention provides an automatic detection method for power safety violation behaviors. Through the knowledge graph-driven multi-modal data construction and dynamic semantic recombination technology, the core problems of insufficient scenario coverage and weak cross-modal semantic association in traditional detection methods are solved. Based on the hierarchical knowledge graph structure (entity layer, event layer, impact layer), combined with causal relationship chains, spatio-temporal relationship chains and compliance relationship chains, diverse scenario description texts are dynamically generated to construct an image-text alignment multi-modal dataset covering the entire power operation scenario; the dual-modal adversarial sample generation technology is introduced to add FGSM / PGD perturbations to images to generate visual adversarial samples, and at the same time, keyword replacement and logical misleading are implemented on text descriptions to generate semantic interference samples, enhancing the robustness of the model to complex interferences. Through the multi-modal feature dynamic fusion model, integrating the knowledge graph path weight optimization mechanism and the joint loss function design, collaborative learning of image features, text semantics and graph structured knowledge is realized, where the path weight calculation integrates node similarity, relationship frequency and expert scoring factors, and multi-modal feature balanced distribution is achieved through Bayesian search and entropy constraint. Based on the model output, a closed-loop feedback mechanism is established. Using the CLIP model to quantify the cross-modal alignment degree of images and texts, triggering the re-annotation of low-confidence samples and the generative completion of missing scenarios, generating realistic-style images through Stable Diffusion and constraining FID≤15, combined with the compliance verification of equipment models and protection specifications by power safety experts, realizing the continuous iterative optimization of the dataset and model parameters.
[0004] Specifically, the following technical solutions are adopted: An automatic detection method for power safety violation behaviors, comprising: Constructing a multi-modal dataset covering the power operation scenario through multi-view image acquisition and knowledge graph-driven dynamic semantic recombination, and the dynamic semantic recombination generates diverse language descriptions based on the hierarchical structure and path weight optimization of the knowledge graph; Expanding the dataset based on the adversarial sample generation technology to generate image perturbation samples and semantic misleading text descriptions; Training an object detection network through a multi-modal feature dynamic fusion model, and the model realizes image-text correlation learning through knowledge graph path weight optimization and joint loss function; Inferring the real-time acquired power operation images based on the trained model, and outputting the detection results of violation behaviors and their semantic descriptions; Triggering a closed-loop feedback through multi-modal similarity evaluation to optimize dataset annotation and model parameters.
[0005] Furthermore, the hierarchical structure of the knowledge graph includes: The entity layer, including entity of personnel, equipment and location; The event layer connects violation behaviors and risk consequences through direct causation, occurrence in, and violation relation types; The impact layer defines the risk level, associated regulatory clauses, and consequence descriptions.
[0006] Furthermore, the dynamic semantic recombination includes: Synonym replacement: Generate diverse description variants based on the power safety term mapping table and enhance semantic coherence through a natural language generation model; Boolean logic generation: Based on the logical relationships in the event layer of the knowledge graph, generate compound risk scenario descriptions by combining multiple violation behaviors; Situational awareness: Dynamically expand time, weather, and equipment information to generate detailed descriptions; Scene description generation: Based on the causal relationship chain and spatio-temporal relationship chain of the knowledge graph, dynamically combine violation behaviors, risk consequences, and context information.
[0007] Furthermore, the natural language generation model is a T5 model fine-tuned on power safety domain texts to generate context-adapted synonym combinations; The compound risk scenario description combines violation behaviors through logical operators AND / OR and associates with the risk consequences in the impact layer of the knowledge graph; The situational awareness dynamically extracts time, location, and equipment information through the spatio-temporal relationship chain of the knowledge graph to generate detailed descriptions.
[0008] Furthermore, the knowledge graph path weight optimization includes: Node vectorization: Generate text feature vectors based on BERT fine-tuned on power safety specification texts and generate structural feature vectors through GNN; Context feature extraction: Generate time, location, and operation status feature vectors through spatio-temporal encoders.
[0009] Furthermore, the comprehensive similarity calculation formula for the path weights is:
[0010] Where: Represents the cosine similarity between two vectors; Are the weight coefficients of text features, structural features, and context features respectively, satisfying ; The weights are dynamically optimized through Bayesian search; v i And v j Respectively represent the text feature vectors of two nodes (such as entities or events in the knowledge graph); s i And s j Respectively represent the structural feature vectors of two nodes; c i And c jRefers to the context feature vector, including information such as time, location, and operation status, generated by the spatio-temporal encoder; The calculation fusion relationship of the path weight combines the frequency factor and the expert score, and the formula is:
[0011] Where is the similarity between nodes and ; is the frequency of the relationship ; is the semantic correlation score between two nodes; the parameter β is a parameter used to adjust the steepness of the curve; is the normalized expert score that maps the expert score to the interval [0, 1] and are weight adjustment functions, f The function is used to adjust the similarity score, g The function is used to adjust the weight according to the relationship frequency; h is a function of semantic correlation, used to measure the semantic association degree between nodes.
[0012] Furthermore, the adversarial sample generation technology includes: Adding perturbation momentum to the original image through the FGSM or PGD algorithm to generate visual adversarial samples; Generating text adversarial samples by replacing keywords or adjusting logical relationships.
[0013] Furthermore, the closed-loop feedback includes: Using the CLIP model to calculate the image-text similarity score, and triggering the relabeling process when it is lower than the threshold; Generating missing scene images through Stable Diffusion, constraining the FID score of the generated images ≤ 15, and verifying compliance through power safety experts.
[0014] And, an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the program, the steps of the above method are implemented.
[0015] A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the above method are implemented.
[0016] Compared with the prior art, the present invention and its preferred solutions have at least the following beneficial effects: Through the knowledge graph-driven multimodal data dynamic reorganization technology, an image-text alignment dataset covering the entire power operation scene is constructed, breaking through the semantic association limitations of traditional single-modal datasets, and significantly improving the scene coverage capability and cross-modal semantic understanding accuracy of complex violations; Robustness optimization: Combining bimodal adversarial sample generation technology and adding domain-adaptive interference to images and texts, the model can effectively enhance its resistance to visual noise (such as lighting changes and device occlusion) and semantic misleading (such as keyword replacement and logical contradictions). Improved detection accuracy: Based on the multimodal feature dynamic fusion mechanism optimized by knowledge graph path weight, the collaborative reasoning of image features, text descriptions and structured knowledge (such as the provisions of the "Electric Power Safety Work Regulations") is realized, and the detection accuracy and decision interpretability of compound violations (such as "not wearing a seat belt and illegal operation") are improved; Closed-loop iteration advantage: Through the closed-loop feedback mechanism of cross-modal similarity evaluation and generative data completion, the dataset annotation quality and model generalization ability are dynamically optimized to solve the problem of model performance degradation caused by data staticization in traditional methods; In addition, a path weight allocation strategy combining Bayesian optimization and entropy constraints is adopted to balance the influence of text, structure and context features on the detection results, avoiding false detection caused by the dominance of a single modality; the deep integration of the expert scoring mechanism and the knowledge graph ensures that the model decision-making complies with the professionalism and authority of power safety regulations. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: Figure 1 This is a technical flow chart of an embodiment of the present invention for realizing automatic detection of power safety violations; Figure 2 is a technical flow chart of the automatic detection step (1) of power safety violations implemented in an embodiment of the present invention; Figure 3 is a technical flow chart of the automatic detection step (2) of power safety violations implemented in an embodiment of the present invention; Figure 4 is a technical flow chart of the automatic detection step (3) of power safety violations implemented in an embodiment of the present invention; Figure 5 is a technical flow chart of the automatic detection step (4) of power safety violations implemented in an embodiment of the present invention; Figure 6 is a technical flow chart of the automatic detection step (5) of power safety violations implemented in an embodiment of the present invention; Figure 7 is a technical flow chart of the automatic detection step (6) of power safety violations implemented in an embodiment of the present invention; Figure 8 It is the dataset architecture diagram for the automatic detection of power safety violation behaviors implemented in the embodiments of the present invention; Figure 9 It is the technical flow chart of step (8) for the automatic detection of power safety violation behaviors implemented in the embodiments of the present invention; Figure 10 It is the schematic diagram of the framework of the BERT model in the embodiments of the present invention; Figure 11 It is the example diagram of the dataset after positive and negative sample pairing in the embodiments of the present invention. Detailed implementation manners
[0018] To make the features and advantages of the present invention more obvious and understandable, specific embodiments are hereinafter given and described in detail as follows: It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0019] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0020] As Figures 1 - 11 shown, the embodiments of the present invention provide an automatic detection method for power safety violation behaviors, and the specific steps are as follows: Step (1) Collect multi-perspective and multi-scenario c images covering power operation personnel and their violation behaviors, and use geometric transformation and image scaling techniques to unify all images to a specified size , specifically: Objective setting: Clearly define the types of power operation scenarios (c) to be collected and the different perspectives ( ) to be captured in each scenario, including but not limited to high-altitude operations, substation maintenance, etc.; Equipment preparation: Select appropriate photography and videography equipment to ensure that high-quality images can be captured from different angles. Consider using multiple devices such as drones, hand-held cameras, or fixed cameras; On-site shooting: Dispatch a professional team to different power operation sites to collect images according to the preset perspectives and scenarios. Ensure that as many scenarios and perspectives as possible are covered to obtain a comprehensive dataset; Simulation environment shooting: In the case where actual operation images cannot be directly obtained, a simulation environment is constructed for shooting to ensure the diversity of the dataset.
[0021] Initial screening and classification: In the initial screening and classification stage, first, quality detection and evaluation are performed on all the collected images. An automated script is used to detect the quality of each image, specifically including blur (preferably measured by calculating the Laplacian variance) and exposure (preferably evaluated using histogram analysis). For blur, if the Laplacian variance of an image is below the set threshold, the image is considered too blurred and not suitable as a training sample; for exposure, by analyzing the brightness histogram of the image, it is judged whether it is within a reasonable range to avoid overexposure or underexposure. In addition, other quality issues need to be checked, such as whether there are obvious noise points or color distortion in the image.
[0022] After completing the quality detection, the remaining high-quality images are classified according to different scenarios (such as high-altitude operations, substation maintenance, etc.) and perspectives (such as top view, bottom view, side view, etc.). Ensure that the images under each scenario and perspective are correctly classified, and a clear file structure is established for subsequent annotation and processing work. This process not only helps to improve the overall quality of the dataset but also lays a solid foundation for the training of the multi-modal object detection model.
[0023] Constructing the image set: Let there be a finite set of perspectives and a finite set of scenarios , then all the multi-perspective and multi-scenario image sets can be represented in the form of a multi-dimensional array or matrix:
[0024] where represents the image under the i-th perspective and the j-th scenario ; Size standardization: Using image processing software (such as OpenCV, PIL, etc.), geometric transformations (rotation, flipping, cropping, etc.) and scaling techniques are applied to each image to adjust all the images to a unified target size . When performing image scaling, note to keep the ratio of the original image unchanged to avoid distortion. If necessary, padding can be added to the image edges to reach the target size; Data augmentation: To increase the diversity of model training, data augmentation operations can be performed on the adjusted images, such as random cropping, color jittering, noise addition, etc., and the processed images are properly saved by category and perspective, and a clear file structure is established for easy access and use in subsequent steps.
[0025] Step (2) Design a series of language descriptions for power violation behaviors , ranging from simple behavior tags to detailed scenario descriptions , ensuring description flexibility, specifically including: Clarify the purpose of generating descriptions and the requirements to be met: Short behavior tags and detailed scenario descriptions are needed.
[0026] Professional term extraction: Collect professional terms related to power operations, including types of violation behaviors (such as "not wearing a safety helmet"), occurrence locations (such as "inside the substation"), involved equipment (such as "high-voltage switch"), risk levels (such as "high risk"), etc.; Use natural language processing techniques to extract high-frequency words: Such as word frequency statistics, TF-IDF analysis, extract high-frequency words from literature and specification documents in the field of power safety; Introduce context-sensitive words: For example, according to different weather conditions (sunny, rainy), time (day, night) or operation status (maintenance, overhaul), dynamically adjust the keywords in the description; Design behavior tags : Design short behavior tag templates to directly describe the core actions of violation behaviors, such as "${type} type of violation behavior", "${type} type of violation behavior occurring at ${location}", "${type} type of violation behavior in ${state}" state", replace the placeholders (such as ${type}, ${location}, ${state}) to generate diverse tags; where ${…} is used to represent placeholders, which is a marking method for template variables, used to refer to specific content or values. When creating specific descriptions or sentences, these placeholders will be replaced by actual data or information.
[0027] Expand behavior tags into scenario descriptions : Expand behavior tags and combine specific scenario information to generate detailed descriptions, such as "At ${location}, a staff member committed a ${type} type of violation behavior, not wearing ${equipmen}, resulting in an increased ${risk} risk."; "${worker} committed a ${type} type of violation behavior when performing tasks at ${location} at ${time}, violating the ${standard} standard." Introduce a natural language generation (NLG) model: Utilize the powerful context understanding ability of pre-trained models to automatically generate more diverse and complex descriptions; Input the existing behavior tags and scenario description templates into these models to generate more diverse descriptions.
[0028] Collect and organize existing behavior tags and scenario description templates to form a standardized dataset; use the above dataset to fine-tune the T5 model so that it can better understand and generate descriptions related to electricity violation behaviors; input the given basic description template into the fine-tuned model to generate various descriptions. For example, input the template: "At ${location}, a staff member committed a ${type} violation behavior, not wearing ${equipment}, resulting in an increased ${risk} risk."; Output example 1: "In the substation, an electrician committed a violation of not wearing a safety helmet, increasing the risk of head injury."; Output example 2: "In the distribution room, technicians operated without wearing protective helmets as required, posing a serious safety hazard." Evaluate and verify the generated descriptions to ensure their accuracy and applicability. Continuously optimize the model parameters and generation strategies based on the feedback.
[0029] Test, verify, and adjust the optimized description template: Test the designed template to evaluate its effectiveness and applicability. The flexibility and accuracy of the template can be verified by simulating different scenarios. Adjust and optimize the template according to the test feedback to ensure that the template can meet the actual application requirements. Confirm that the template reaches the expected goal and complete the design.
[0030] Step (3) Generate descriptions based on the knowledge graph. By defining key semantic roles, constructing a structured knowledge graph, and extracting path weights, select appropriate description templates to generate diverse descriptions. Specifically: Define the key semantic roles in violation behaviors: including the subject (Who), action (What), object (Where / Which), status (How), etc. Specific examples are shown in Table 1 below: Table 1 Examples of key semantic roles in violation behaviors
[0031] According to different combinations of semantic roles, dynamically generate language descriptions. For example, the subject is a staff member; the action is not wearing; the object is a safety helmet; the status is not in accordance with the regulations → the description is "The staff member did not wear a safety helmet as required." Construct a structured knowledge graph: Hierarchical modeling: The knowledge graph is divided into three layers: the entity layer, the event layer, and the impact layer.
[0032] Entity Layer: Personnel: including staff members, maintenance workers, technicians, etc.
[0033] Equipment: such as high-voltage switches, circuit breakers, insulating gloves, protective clothing, etc.
[0034] Location: For example, substations, distribution rooms, power facilities, etc.
[0035] Event Layer: Violations: failure to wear a safety helmet, illegal operation of high-voltage switches, failure to wear a seat belt, etc.
[0036] Compliance operations: Wear a safety helmet correctly, operate high-voltage switches according to regulations, use insulating tools correctly, etc.
[0037] Impact Layer: Risk: high risk, medium risk, low risk, etc.
[0038] Consequences: risk of electric shock, falling from height, fire hazard, etc.
[0039] Regulatory provisions: violation of safety standards, compliance with safety regulations, exemption clauses, etc.
[0040] Expanded relationship types: Based on the original simple relationships, multiple relationship types are added to more accurately describe complex causal chains and spatiotemporal relationships.
[0041] Causal Relations: Directly Leads To: Indicates that one event directly leads to the occurrence of another event.
[0042] Indirectly Triggers: Indicates that one event indirectly triggers another event through a series of intermediate events.
[0043] Temporal and Spatial Relations: Occurs At: Indicates the place where the event occurs.
[0044] Lasts Until: Indicates the duration of an event.
[0045] Compliance Relations: Violates: Indicates that an action violates a specific safety standard or regulation.
[0046] Complies With: Indicates that an action complies with specific safety standards or regulations.
[0047] Exempts From: Indicates that certain circumstances may exempt you from complying with specific safety standards or regulations.
[0048] Path weight calculation: Node vectorization and similarity calculation: Use a pre-trained word vector model (such as BERT) to generate vector representations for each node. To improve domain adaptability, when fine-tuning the BERT model, add power safety texts (such as the full text of the "Power Safety Work Regulations" and accident reports), as Figure 10 shown.
[0049] Assume that each node and have text descriptions of and respectively, then their vector representations are and respectively.
[0050]
[0051]
[0052] v i and v j represent the text feature vectors of two nodes (such as entities or events in a knowledge graph) respectively, and BERT( ) is used to generate text feature vectors. These vectors are usually generated by a pre-trained word vector model (such as BERT) and are fine-tuned for domain adaptability to enhance the understanding ability in the power safety domain.
[0053] Utilize the structural information of nodes in the knowledge graph to enhance similarity calculation. Specifically, the structural features of nodes can be extracted through a graph neural network (GNN). Assume that and are the structural feature vectors of nodes and respectively.
[0054]
[0055]
[0056] s i and s j represent the structural feature vectors of two nodes, and GNN( ) is used to extract the structural features of nodes from the knowledge graph. This is usually extracted through a graph neural network (GNN) from the knowledge graph, reflecting the position of the node in the network and its relationship with other nodes.
[0057] Considering the context information of nodes (such as time, location, etc.), similarity calculation can be further enhanced. Assume that and are nodes and If there is context information, their vector representations are respectively and .
[0058]
[0059]
[0060] c i and c j refer to the context feature vectors, which include information such as time, location, and operation status, and are generated by the spatio-temporal encoder. ContextEncoder( ) is responsible for encoding context information such as time, location, and operation status into vector representations.
[0061] Combine the above three features (text, structure, context) to calculate the comprehensive similarity. The specific formula is as follows:
[0062] Where: represents the cosine similarity between two vectors; are the weight coefficients of text features, structure features, and context features respectively, satisfying . Specifically, according to the knowledge of domain experts, the importance of each feature is initially evaluated and the initial weights are set. Using the Bayesian optimization method, the weight coefficients are dynamically adjusted during the training process. Design a feedback-based learning and correction mechanism to update the model parameters in real time to adapt to new data, construct a joint loss function, and optimize the interaction between multiple features at the same time. Since the weight may exceed the reasonable range due to no constraints during the Bayesian optimization process (such as →1, completely ignoring structure and context features), force each weight coefficient ≥ 0.2 to ensure that multi-modal features participate evenly. And add a weight entropy penalty term to the objective function of Bayesian optimization to encourage uniform weight distribution: , where represents the regularization loss function; λ is a hyperparameter called the regularization strength coefficient; k ∈ {t, s, c}, where k is an index variable representing an element in the set {t, s, c}, and t, s, c represent the weights of time features, structure features in the knowledge graph, and context features respectively; for each k, w k represents the corresponding weight value. These weights usually come from the parameters learned during the model training process, and they reflect the importance or contribution degree of different types of features.
[0063] Weight formula optimization: The weights can be calculated by the following formula, which combines node similarity, relationship frequency, expert scores, and semantic relevance between nodes:
[0064] where is the similarity between nodes and ; is the frequency of the relationship , that is, traversing the entire knowledge graph to count the number of times a specific relationship appears as an edge in the graph; is the semantic correlation score between two nodes, used to measure their semantic association degree; the parameter β is a parameter used to adjust the steepness of the curve; is the normalized expert score that maps the expert score to the interval [0, 1]; and are weight adjustment functions (such as linear or non - linear functions), f The function is used to adjust the similarity score. It is preferred to use the Sigmoid function to amplify the difference, making the high similarity score more prominent, g The function is used to adjust the weight according to the relationship frequency, and logarithmic normalization is preferably adopted; h is the function of semantic correlation used to measure the semantic association degree between nodes.
[0065]
[0066] where expert_score is the expert score, and min_score and max_score correspond to the minimum and maximum values in the expert score respectively.
[0067]
[0068] is the input similarity score; α is a parameter that controls the steepness of the curve. A larger α value will make the function change more violently, thus being more sensitive to the change of the similarity score.
[0069]
[0070] y is the input relationship frequency; is the maximum frequency value among all considered relationships.
[0071]
[0072] where is the semantic correlation score between nodes and , is the parameter that controls the steepness of the curve. The semantic correlation score can be obtained by calculating the semantic distance between the descriptions of two nodes through a pre - trained language model (such as BERT).
[0073] Template matching and generation: Template design: Define templates, such as Template 1: At ${location}, ${who} caused ${consequence} (violating ${standard}) due to ${action} ${object}; Template 2: When ${worker} was operating at ${location} at ${time}, failure to ${action} as required increased the risk of ${consequence}; Synonym replacement: Increased risk → Pose... hidden dangers, increase... threats; Not wearing → Not donning, not properly wearing, not wearing as required; Multi-path fusion: Attention mechanism: Calculate the normalized weight of each path Where 𝑊 is the sum of the weights of all paths. Weighted average the descriptions of each path.
[0074] Conflict resolution rule: Select the consequence corresponding to the path with the highest weight as the final result.
[0075] Post-processing verification: Fluency correction: Use a language model (such as GPT-2) to correct the fluency of the generated text.
[0076] Dependency syntactic analysis: Ensure that the subject-verb-object structure is correct.
[0077] Step (4) Semantic recombination and detail supplementation. By performing synonym replacement on the key terms in the description, adjusting the sentence structure to reorganize the information order, and adding more context information to make the description more specific and detailed. Specifically: Synonym replacement: Use a term mapping table or a pre-trained language model to perform synonym replacement to generate new description variants. The specific steps are as follows: Construct a term mapping table specific to the power field, including but not limited to the content in Table 2: Table 2 Special terms in the power field and their synonyms / related expressions Term Synonym / Related Expression Safety helmet Helmet, protective helmet, work helmet Insulating gloves Protective gloves, work gloves, insulating equipment Protective clothing Work clothes, insulating clothing, protective equipment Substation Distribution room, power facilities, high - voltage station, power supply station Electrician Maintenance worker, technician, staff member, maintenance personnel Daily maintenance work Equipment inspection, routine maintenance, daily inspection, maintenance operation Not worn Not put on, not worn correctly, not worn as required Not fastened Not fastened correctly, not fixed, not tied tightly High - altitude operation Working at height, high - altitude construction, working at heights Electric shock risk Electrical hazard, current injury, electric shock risk High risk Extremely high risk, major hidden danger, serious threat Medium risk Higher risk, general hidden danger, medium threat Low risk Minor hidden danger, low - degree threat, controllable risk Violation of operation Improper operation, wrong operation, violation of regulations Insulating tool Protective tool, safety appliance, insulating equipment Emergency passage Evacuation route, escape route, emergency exit Prohibited entry area Dangerous area, restricted area, isolation area High - voltage switch Circuit breaker, distribution switch, high - voltage equipment Personal protective equipment (PPE) Protective supplies, safety equipment, labor protection supplies Falling from height Falling accident, falling from height, falling risk Fire risk Combustion hidden danger, fire threat, thermal runaway risk Unauthorized Unqualified, uncertified, unpermitted Maintenance record Maintenance log, operation record, equipment ledger Emergency plan Emergency handling plan, emergency response plan, accident response measures For the key terms in the input description, replace them with synonyms one by one. For example, for the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One of the electricians is not wearing a safety helmet, and another electrician is not wearing a safety belt." Replacement rules: Safety helmet → Helmet, Substation → Switchgear room, Electrician → Maintenance worker; Generated description after replacement: Let the input description be D, the term set be , the synonym set be , s ikDenote the k-th synonym of the i-th term, then the replaced description can be expressed as: , where Replace is the replacement function, and D is the description, and t i is the i-th term.
[0078] For example, the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One electrician is not wearing a safety helmet, and another electrician is not wearing a seat belt." Replacement result 1: "In a certain power distribution room, multiple maintenance workers are performing daily maintenance work. One maintenance worker is not wearing a helmet, and another maintenance worker has not properly fastened the protective equipment." Replacement result 2: "In a certain power facility, multiple technicians are performing equipment inspections. One technician is not wearing a protective helmet, and another technician is not wearing work clothes as required." Situational awareness: Extract context information from the input description, such as weather conditions, time, operation status, etc.; According to the extracted context information, dynamically adjust the keywords and expressions in the description.
[0079] For example, the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One electrician is not wearing a safety helmet, and another electrician is not wearing a seat belt."; The dynamically adjusted description: "On a sunny day, in the substation, multiple electricians are performing daily maintenance work. Due to not wearing a safety helmet and not wearing a seat belt, one electrician may face the risk of head injury, and another electrician has the potential hazard of falling from a height." Semantic reorganization: Generate descriptions with the same semantics but different expressions by adjusting the sentence structure or reorganizing the information order.
[0080] Adjust the sentence structure or reorganize the information order: Subject-verb-object adjustment, that is, change the order of the subject, verb, and object of the sentence; Clause nesting, that is, nest some information into a clause; Parallelism and progression, that is, use parallel sentences or progressive sentences to enhance the logic of the description.
[0081] Let the input description be D, and the sentence structure be , and S is a set containing three elements: Subject, Verb, and Object. Then the reorganized description can be expressed as , where Reorganize is the semantic reorganization function.
[0082] As the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One electrician is not wearing a safety helmet, and another electrician is not wearing a safety belt.", Reorganization result 1: "Due to not wearing a safety helmet and not wearing a safety belt, multiple electricians working in a certain substation are at risk of falling from a height.", Reorganization result 2: "In a certain substation, although multiple electricians are performing daily maintenance work, one electrician is not wearing a helmet correctly, and another electrician is not wearing a safety belt as required." Add detail supplement: By introducing more context information or expanding the description content, make the generated description more specific and detailed.
[0083] Introduce more context information or expand the description content: Supplement the risk level, that is, clarify the specific risks that the violation behavior may cause; time and location, that is, add the time and specific location of the event; equipment information, that is, supplement the name or model of the equipment involved.
[0084] Let the input description be D, and the supplementary information be , C is a set containing four elements: Risk, Time, Location, and Equipment. Then the supplemented description can be expressed as: , where Expand is the supplementary function.
[0085] As the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One electrician is not wearing a safety helmet, and another electrician is not wearing a safety belt.", Supplementary result 1: "In a certain substation (number #123), multiple electricians are performing daily maintenance work on high-voltage lines at 9 am. One electrician is not wearing a protective helmet, which may cause head injuries; another electrician is not wearing a safety belt as required, posing a risk of falling from a height.", Supplementary result 2: "In a certain substation, multiple electricians are performing maintenance work on transformer equipment. One electrician is not wearing an insulating helmet correctly, which may cause an electric shock accident; another electrician is not wearing protective equipment as required, increasing the overall safety hazard of the operation." Generate a composite description by combining Boolean logic: According to the combination relationship of multiple violation behaviors, generate a composite description containing multiple violation behaviors.
[0086] Let the set of violation behaviors be , and the logical relationship be L. Then the composite description can be expressed as: , where Combine is the combination function.
[0087] If the set of violation behaviors: B = {not wearing a safety helmet, not wearing a safety belt, not wearing protective clothing}, logical relationship: , Output description: "In a certain substation, multiple electricians are carrying out daily maintenance work. One of the electricians is neither wearing a safety helmet nor properly fastening the safety belt. At the same time, another electrician is not wearing protective clothing as required, resulting in multiple safety hazards in the overall operation." Step (5) details the illegal behaviors in each image, including bounding box localization, instance segmentation mask, and behavior status marking, and supplements missing data by combining text-to-image generation technology. Specifically: Bounding box localization: Use professional image annotation tools (such as LabelImg or CVAT) to draw accurate bounding boxes for each instance of illegal behavior and record their coordinate positions ( , , , ), where x min , y min represent the abscissa and ordinate of the lower left corner of the rectangle respectively, and x max , y max represent the abscissa and ordinate of the upper right corner of the rectangle respectively. At the same time, associate the bounding box with the behavior label generated in step S2, such as "not wearing a safety helmet" or "operating the high-voltage switch illegally"; Instance segmentation mask: Use manual segmentation tools (such as LabelMe) or combine pre-trained instance segmentation models (such as Mask R-CNN) to generate pixel-level segmentation masks to ensure that every detail of the target area is covered. Correlate the mask with a more detailed scene description , such as "In the substation, the staff are not wearing protective equipment as required and are performing dangerous operations"; Add detailed status information to each instance of illegal behavior: including but not limited to timestamp, location information, number of involved personnel, etc.
[0088] Language-vision association: Ensure the consistency and accuracy of annotation by quantifying the correlation between images and texts. For example, use a pre-trained CLIP model to extract image and text features and calculate the similarity. If the similarity is lower than a threshold (such as 0.75), determine that the image and text do not match and trigger the re-annotation process; Check for missing pictures: After completing the above annotation and language-vision association operations, if it is found that some descriptions do not have corresponding pictures, pictures can be generated through text-to-image synthesis technology.
[0089] Preprocess the input description (description without corresponding images): Use a BERT-based NER model to identify entities (locations, personnel, actions, objects) in the description; dynamically add details in combination with an external knowledge base (power safety documents): e.g., "The staff member did not wear a safety helmet" → enhanced to: "An electrician wearing work clothes did not wear a safety helmet when conducting equipment inspections inside the substation on a sunny day." Generate images based on the text: Input the enhanced description into Stable Diffusion to generate images: Select a resolution of 1024×1024 (high detail is required for power scenarios); realistic style (artistic style disabled); generate 3 - 5 candidate images for screening.
[0090] Further optimize the generated images: Use OpenCV to adjust brightness / contrast to ensure consistency with the real image distribution, and screen the image with the highest text similarity based on CLIP.
[0091] Quality control: Calculate the FID score (Frechet Inception Distance), requiring that the FID between the generated image and the real image ≤ 15; verify the semantic accuracy of the generated image (such as safety helmet model, scene compliance) by power safety experts.
[0092] Step (6) Create positive and negative sample pairs based on the language description , where the positive sample pair corresponds to the accurately described behavior, while the negative sample contains behaviors that do not match or are irrelevant to the description. Specifically: Select the matching description: According to the type and detailed information of the violation behavior marked in the image, select the most appropriate language description as the positive sample description. For example, if the image shows an electrician not wearing a safety helmet, select a description like "Inside the substation, an electrician is conducting maintenance work without wearing a safety helmet as required"; Generate positive sample pairs: Pair the image with its corresponding correct description to form positive sample pairs (image,description), where image refers to the image data containing a specific violation behavior; description is the text description that precisely matches the above image and accurately reflects the violation behavior in the image. Ensure that each pair is precisely matched to avoid any inconsistent situations; Generate non - matching descriptions: Generate non - matching descriptions through random replacement, deliberate errors, and scene dislocation: Randomly select descriptions from the language description library that are irrelevant to the violation behavior of the current image, or deliberately select descriptions that are opposite to the type of violation behavior in the image (e.g., if the image shows "not wearing a safety helmet", the description is "wearing a safety helmet correctly"), and select descriptions for another scene (e.g., if the image shows working at height, the description is "operating inside the substation"); Introduction of adversarial sample generation: Image adversarial samples: Use adversarial attack algorithms (such as FGSM, PGD, etc.) to add tiny perturbations to an image, making the image look almost the same, but causing the model to produce incorrect classifications or descriptions. Generate images with tiny perturbations using an adversarial attack algorithm and assign them the same description as the original image. Such samples can test the sensitivity of the model to subtle changes.
[0093] Description adversarial samples: By changing some keywords or phrases in the description to make it semantically seem reasonable but actually not match the image. For example, keep most of the description unchanged but make misleading modifications in the key parts. Make minor semantic adjustments to the original description to generate a description that seems reasonable but actually does not match. For example, if the image shows "not wearing a safety helmet", the description is "Although not wearing a safety helmet, other protective equipment is complete".
[0094] Generate negative sample pairs: Pair an image with a description that does not match it to form negative sample pairs (image, incorrect_description). Ensure that each pair does not match to train the model to distinguish between correct and incorrect descriptions; where image refers to the image data containing a specific violation behavior; incorrect_description is a deliberately mismatched description, that is, the content it describes does not match the actual situation in the image.
[0095] Image adversarial sample pairs: Pair the image with tiny perturbations with its original description to form adversarial sample pairs (adversarial_image, original_description), where adversarial_image is the image generated by adding tiny perturbations to the original image using an adversarial attack algorithm (such as FGSM, PGD, etc.); original_description is the original correct description corresponding to the adversarial sample image, and even if the image has been perturbed, its description should remain unchanged.
[0096] Description adversarial sample pairs: Pair the original image with the description after semantic adversarial adjustment to form adversarial sample pairs (original_image, adversarial_description), where original_image refers to the real-scene image without any modification or perturbation, which accurately reflects a specific power operation scene and possible violations; adversarial_description is the description generated through specific processing to make it not match the content of the original image or contain misleading information.
[0097] To ensure the quality of positive and negative sample pairs, rigorous review and validation are performed, including cross-checking, expert review, and establishing a feedback mechanism to continuously improve the accuracy and consistency of sample pairs.
[0098] Step (7) Integrate the results of the previous steps to build a comprehensive data set of power safety violations covering multiple scenarios and complex language descriptions , specifically: Summarize and organize all the data generated in the previous steps, including images, annotation information, language descriptions and their corresponding positive and negative sample pairs: Image data: all labeled images of power violations, including bounding box locations, instance segmentation masks, and behavior status labels; Language description: covers a variety of language description templates from short behavior labels to detailed scenario descriptions; Positive and negative sample pairs: ensure that each image has a corresponding positive sample pair (accurately described behavior) and negative sample pair (inconsistent or irrelevant behavior); e.g. Figure 11 As shown, the positive sample description P given by the corresponding picture I There are: "Workers wearing blue helmets", "Workplaces with safety fences", "Workers next to a white truck are on the phone"; negative sample description N I There are: "There are no workers wearing safety helmets", "There are no illegal operations in the picture"; Step (8) trains a multimodal detection model by quantifying the correlation between images and texts and continuously optimizes the dataset labels to improve the accuracy and robustness of the model. Specifically: Dataset division: divide the dataset into training set, validation set and test set to ensure balanced distribution; Model selection: Select ViLBERT, a model architecture suitable for multimodal tasks, for training. This model can process image and text inputs simultaneously and learn the correlation between them. Loss function design: Design a suitable loss function (such as contrast loss, cross entropy loss, etc.) to guide model training. The following is the design of the specific loss function: Contrastive Loss: The goal of contrastive loss is to bring positive pairs closer together while pushing negative pairs further apart. Assumptions: The feature vector representing the image x; The feature vector representing the text description y; represents the Euclidean distance between the two; m>0 is a predefined threshold used to control the distance between negative sample pairs. The contrast loss formula is:
[0099] in, N is the total number of samples; is a label indicating whether the sample pair is a positive sample pair ( ) or a negative sample pair ( ); Cross-Entropy Loss: Cross-Entropy Loss is used for supervised classification tasks to calculate the difference between the model's predicted distribution and the true distribution. Assume that: is the class probability distribution predicted by the model; is the one-hot encoded distribution of the true label. The Cross-Entropy Loss formula is:
[0100] where N is the total number of samples; C is the total number of classes; is the true label of sample ; is the probability distribution predicted by the model.
[0101] Combined Loss Function: Combine the contrastive loss and the cross-entropy loss to form a combined loss function. Assume that: and λ 2 are two hyperparameters used to balance the importance of the two losses. The combined loss function formula is: . If more attention is paid to the matching relationship between the image and the text, increase ; if more attention is paid to the accuracy of the classification task, increase .
[0102] Feature Extraction: Use the visual encoder of ViLBERT to extract feature vectors from images; use the language encoder of ViLBERT to extract feature vectors from text descriptions.
[0103] Calculate the relevance between the image and the text to evaluate their degree of correlation: Use cosine similarity, dot product similarity, or other similarity measurement methods to quantify the relevance between the image and text features. For example, for each image-text pair, calculate the cosine similarity score between their feature vectors.
[0104] Optimize Model Parameters: Through the backpropagation algorithm, adjust the model parameters according to the gradient of the contrastive loss function to maximize the similarity score between positive sample pairs and minimize the similarity score between negative sample pairs.
[0105] Regularly Evaluate Model Performance: Regularly evaluate the model performance (such as accuracy, recall, F1 score), and adjust hyperparameters, increase training data, or improve the model architecture according to the evaluation results. For example, if the model performs poorly on certain specific types of violation behaviors, the model effect can be improved by increasing the data of these types.
[0106] Model Deployment: Deploy the trained multi-modal model to the actual application scenario for real-time identification of electricity violation behaviors. This model can make full use of the design of the previous series of high-quality data sets to achieve accurate identification and efficient management of violation behaviors in electricity operations.
[0107] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application-Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically used to load and execute one or more instructions in the computer storage medium to implement the above method.
[0108] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program executes the above method when run by a processor. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.
[0109] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
[0111] The present invention is not limited to the above best implementation manner. Anyone can obtain various other forms of automatic detection methods for power safety violation behaviors under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage scope of the present invention.
Claims
1. A method for automatically detecting power safety violations, characterized in that: A multimodal dataset covering power operation scenarios is constructed through multi-view image acquisition and knowledge graph-driven dynamic semantic reorganization. The dynamic semantic reorganization generates diversified language descriptions based on the hierarchical structure and path weight optimization of the knowledge graph. Expand the dataset based on adversarial sample generation technology to generate image perturbation samples and semantically misleading text descriptions; The object detection network is trained through a multimodal feature dynamic fusion model, which realizes image-text association learning through knowledge graph path weight optimization and joint loss function; Based on the trained model, the real-time collected power operation images are inferred to output the violation behavior detection results and their semantic descriptions; The closed-loop feedback is triggered by multimodal similarity evaluation to optimize dataset annotation and model parameters.
2. The method for automatically detecting power safety violations according to claim 1, characterized in that: The hierarchical structure of the knowledge graph includes: The physical layer includes personnel, equipment, and location entities; Event layer, connecting illegal behaviors and risk consequences through direct cause, occur in and violation relationship types; Impact layer, defines the risk level, associated regulatory provisions and consequence description.
3. The automatic detection method for power safety violations according to claim 1 is characterized by: The dynamic semantic reorganization includes: Synonym replacement: Generate diverse description variants based on the power safety terminology mapping table and enhance semantic coherence through a natural language generation model; Boolean logic generation: Based on the logical relationship of the knowledge graph event layer, a composite risk scenario description is generated by combining multiple violations; Context awareness: dynamically expand time, weather and device information to generate detailed descriptions; Scenario description generation: Based on the causal relationship chain and spatiotemporal relationship chain of the knowledge graph, the violation behavior, risk consequences and contextual information are dynamically combined.
4. The automatic detection method for power safety violations according to claim 3 is characterized by: The natural language generation model is a T5 model fine-tuned based on text in the field of power safety, which generates context-adapted synonym combinations; The composite risk scenario describes the risk consequences of combining illegal behaviors through logical operators AND / OR and associating the risk consequences of the knowledge graph impact layer; The context awareness dynamically extracts time, location and equipment information through the spatiotemporal relationship chain of the knowledge graph to generate a detailed description.
5. The automatic detection method for power safety violations according to claim 1 is characterized by: The knowledge graph path weight optimization includes: Node vectorization: Generate text feature vectors based on BERT fine-tuned on the power safety specification text, and generate structural feature vectors through GNN; Context feature extraction: Generate time, location and job status feature vectors through the spatiotemporal encoder.
6. The automatic detection method for power safety violations according to claim 5 is characterized by: The comprehensive similarity calculation formula of the path weight is: in: Represents the cosine similarity between two vectors; are the weight coefficients of text features, structural features and context features, satisfying ;Weights are dynamically optimized through Bayesian search;v i and v j Represents the text feature vectors of two nodes respectively; s i and j Represents the structural feature vectors of two nodes respectively; c i and c j refers to the context feature vector, generated by the spatiotemporal encoder; The calculation of the path weight integrates the relationship frequency factor and the expert score, and the formula is: in: Is a node and similarity; It's a relationship Frequency is the semantic relevance score between two nodes; parameter β is a parameter used to adjust the steepness of the curve; It is the normalized expert score that maps the expert score to the interval of 0 and 1; and is the weight adjustment function, f The function is used to adjust the similarity score. g The function is used to adjust the weights according to the frequency of the relationship; h It is a function of semantic relevance, which is used to measure the semantic association between nodes.
7. The automatic detection method for power safety violations according to claim 1 is characterized by: The adversarial sample generation technology includes: Add a slight perturbation to the original image through the FGSM or PGD algorithm to generate a visual adversarial sample; Generate text adversarial samples by replacing keywords or adjusting logical relationships.
8. The method for automatically detecting power safety violations according to claim 1, characterized in that: The closed-loop feedback includes: The CLIP model is used to calculate the image-text similarity score, and the re-annotation process is triggered when it is lower than the threshold; The missing scene images are generated through Stable Diffusion, the FID score of the generated images is constrained to be ≤15, and the compliance is verified by power safety experts.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Information pushing method and system fusing knowledge graph structure and path semantics
CN114265986A
Power violation semantic generation method based on scene graph technology
CN117292379A
Method for batch identification and automatic classification and archiving of picture content
CN118799619A
Method and device for determining operation state of ultra-high voltage isolation switch and storage medium
CN119089400A
Constrained large model patent atlas construction method
CN119415707A
Cited By
Transformer substation peccancy detection method based on real-time open vocabulary detection
CN120510571A
Marketing video auditing method based on AI
CN120583273A
Intelligent analysis method, device and equipment for power violation operation
CN121527700A
Intelligent analysis method, device and equipment for power violation operation
CN121527700B
Operation violation behavior online identification method based on multi-source data alignment
CN122112797A