Automatic Detection Method for Electric Power Safety Violation Behaviors
Through knowledge graph-driven multimodal data construction and adversarial sample generation technology, the problems of insufficient scenario coverage and weak cross-modal semantic correlation in power safety detection are solved, and high-precision violation detection and interpretability decisions for complex power operation environments are realized.
Patent Information
- Application Number
- CN202510577571.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing power safety detection methods are difficult to adapt in complex power operation environments, lack the recognition accuracy of multi-view occlusion, dynamic lighting changes and subtle violations, and the data set lacks multi-modal labeling information, and the semantic relationship between images and text descriptions is weak, resulting in limited understanding and reasoning capabilities of the model during actual deployment.
Through knowledge graph-driven multimodal data construction and dynamic semantic recombination technology, a diverse image-text alignment data set is generated, adversarial sample generation technology is introduced, and the multimodal feature dynamic fusion model and joint loss function is combined to realize image-text correlation learning, and the data set and model parameters are optimized through a closed-loop feedback mechanism.
It significantly improves the scenario coverage ability and cross-modal semantic understanding accuracy of complex violations, enhances the robustness of the model to complex interference, improves detection accuracy and decision interpretability, and solves the model performance attenuation problem caused by data staticization in traditional methods.
Smart Images

Figure CN120088864B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and deep learning multimodality, etc., and in particular to an automatic detection method for power safety violation behaviors. Background Art
[0002] With the continuous improvement of the digital upgrade of the power industry and the demand for intelligent operation and maintenance, the safety monitoring of power operations has gradually shifted from traditional manual inspections to intelligent technology-driven. Although the object detection technology based on computer vision has achieved certain application results in industrial scenarios, its practice in the field of power safety still faces significant challenges: existing detection methods mostly rely on general scenario training models and are difficult to adapt to the complex environment and refined specification requirements unique to power operations. For example, in scenarios such as the detection of the wearing of high-altitude operation protection equipment and the identification of the compliance of high-voltage equipment operation, traditional algorithms have insufficient recognition accuracy for multi-view occlusion, dynamic light changes, and subtle violation behaviors, and lack in-depth integration of professional domain knowledge such as the "Power Safety Work Regulations". At the same time, there are structural defects at the data level. Existing data sets generally lack multi-modal annotation information covering the entire power scenario, and the semantic association between images and text descriptions is weak, making it difficult to support the cross-modal association learning ability required by large models, resulting in limited understanding and reasoning ability of the model for complex violation scenarios (such as "not wearing a seat belt and not setting up an insulating partition") during actual deployment. This dual lack of data and knowledge seriously restricts the reliability and interpretability of intelligent detection systems in complex power operation environments. Summary of the Invention
[0003] In view of the defects and deficiencies existing in the prior art, the present invention provides an automatic detection method for power safety violation behaviors, which solves the core problems of insufficient scene coverage and weak cross-modal semantic association in traditional detection methods through knowledge graph-driven multi-modal data construction and dynamic semantic recombination technology. Based on a hierarchical knowledge graph structure (entity layer, event layer, impact layer), combined with causal relationship chains, spatio-temporal relationship chains, and compliance relationship chains, diverse scene description texts are dynamically generated to construct an image-text alignment multi-modal dataset covering the entire power operation scenario; the dual-modal adversarial sample generation technology is introduced to add FGSM / PGD perturbations to images to generate visual adversarial samples, and at the same time, keyword replacement and logical misleading are implemented on the text descriptions to generate semantic interference samples, enhancing the robustness of the model to complex interferences. Through a multi-modal feature dynamic fusion model, integrating the knowledge graph path weight optimization mechanism and the joint loss function design, collaborative learning of image features, text semantics, and graph structured knowledge is achieved, where the path weight calculation integrates node similarity, relationship frequency, and expert scoring factors, and multi-modal feature balanced distribution is realized through Bayesian search and entropy constraint. Based on the model output, a closed-loop feedback mechanism is established, using the CLIP model to quantify the cross-modal alignment degree of images and texts, triggering the re-annotation of low-confidence samples and the generative completion of missing scenes, generating realistic-style images through Stable Diffusion and constraining FID≤15, and combining the compliance verification of equipment models and protection specifications by power safety experts to realize the continuous iterative optimization of the dataset and model parameters.
[0004] Specifically, the following technical solutions are adopted:
[0005] An automatic detection method for power safety violation behaviors, comprising:
[0006] Construct a multi-modal dataset covering the power operation scenario through multi-view image acquisition and knowledge graph-driven dynamic semantic recombination, and the dynamic semantic recombination generates diverse language descriptions based on the hierarchical structure of the knowledge graph and path weight optimization;
[0007] Expand the dataset based on the adversarial sample generation technology to generate image perturbation samples and semantic misleading text descriptions;
[0008] Train an object detection network through a multi-modal feature dynamic fusion model, and the model realizes image-text correlation learning through knowledge graph path weight optimization and joint loss function;
[0009] Infer the power operation images collected in real time based on the trained model, and output the detection results of violation behaviors and their semantic descriptions;
[0010] Trigger a closed-loop feedback through multi-modal similarity evaluation to optimize the dataset annotation and model parameters.
[0011] Furthermore, the hierarchical structure of the knowledge graph includes:
[0012] The entity layer, including entity of personnel, equipment and location;
[0013] The event layer, connecting violation behaviors and risk consequences through relationship types of directly causing, occurring in and violating;
[0014] The impact layer, defining risk levels, associated regulatory clauses and consequence descriptions.
[0015] Furthermore, the dynamic semantic recombination includes:
[0016] Synonym replacement: generating diverse description variants based on the power safety term mapping table and enhancing semantic coherence through a natural language generation model;
[0017] Boolean logic generation: generating compound risk scenario descriptions by combining multiple violation behaviors based on the logical relationships in the event layer of the knowledge graph;
[0018] Context awareness: dynamically expanding time, weather and equipment information to generate detailed descriptions;
[0019] Scene description generation: dynamically combining violation behaviors, risk consequences and context information based on the causal relationship chain and spatio-temporal relationship chain of the knowledge graph.
[0020] Furthermore, the natural language generation model is a T5 model fine-tuned based on texts in the power safety field, generating context-adapted synonym combinations;
[0021] The compound risk scenario description combines violation behaviors through logical operators AND / OR and associates with the risk consequences in the impact layer of the knowledge graph;
[0022] The context awareness dynamically extracts time, location and equipment information through the spatio-temporal relationship chain of the knowledge graph to generate detailed descriptions.
[0023] Furthermore, the knowledge graph path weight optimization includes:
[0024] Node vectorization: generating text feature vectors based on BERT fine-tuned through power safety specification texts and generating structural feature vectors through GNN;
[0025] Context feature extraction: generating time, location and operation status feature vectors through a spatio-temporal encoder.
[0026] Furthermore, the comprehensive similarity calculation formula for the path weight is:
[0027]
[0028] Where: Represents the cosine similarity between two vectors; are the weight coefficients of text features, structural features, and context features respectively, satisfying ; The weights are dynamically optimized through Bayesian search; v i and v j represent the text feature vectors of two nodes (such as entities or events in a knowledge graph) respectively; s i and s j represent the structural feature vectors of two nodes respectively; c i and c j refer to the context feature vectors, including information such as time, location, and operation status, and are generated by a spatio-temporal encoder;
[0029] The calculation of the path weight fuses the relationship frequency factor and the expert score, and the formula is:
[0030]
[0031] where is the similarity between nodes and ; is the frequency of the relationship ; is the semantic correlation score between two nodes; The parameter β is a parameter used to adjust the steepness of the curve; is the normalized expert score that maps the expert score to the interval [0, 1]; and are weight adjustment functions. The f function is used to adjust the similarity score, and the g function is used to adjust the weight according to the relationship frequency; h is a function of semantic correlation, which is used to measure the semantic association degree between nodes.
[0032] Furthermore, the adversarial sample generation technology includes:
[0033] Adding perturbation momentum to the original image through the FGSM or PGD algorithm to generate visual adversarial samples;
[0034] Generating text adversarial samples by replacing keywords or adjusting logical relationships.
[0035] Furthermore, the closed-loop feedback includes:
[0036] Using the CLIP model to calculate the image-text similarity score, and triggering the re-annotation process when it is lower than the threshold;
[0037] Generating missing scene images through Stable Diffusion, constraining the FID score of the generated images ≤ 15, and verifying compliance through power safety experts.
[0038] And, an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the steps of the above method when executing the program.
[0039] A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that the computer program implements the steps of the above method when executed by a processor.
[0040] Compared with the prior art, the present invention and its preferred embodiments have at least the following beneficial effects:
[0041] Through the knowledge graph-driven multimodal data dynamic reorganization technology, an image-text alignment dataset covering the entire power operation scene is constructed, breaking through the semantic association limitations of traditional single-modal datasets, and significantly improving the scene coverage capability and cross-modal semantic understanding accuracy of complex violations;
[0042] Robustness optimization: Combining bimodal adversarial sample generation technology and adding domain-adaptive interference to images and texts, the model can effectively enhance its resistance to visual noise (such as lighting changes and device occlusion) and semantic misleading (such as keyword replacement and logical contradictions).
[0043] Improved detection accuracy: Based on the multimodal feature dynamic fusion mechanism optimized by knowledge graph path weight, the collaborative reasoning of image features, text descriptions and structured knowledge (such as the provisions of the "Electric Power Safety Work Regulations") is realized, and the detection accuracy and decision interpretability of compound violations (such as "not wearing a seat belt and illegal operation") are improved;
[0044] Closed-loop iteration advantage: Through the closed-loop feedback mechanism of cross-modal similarity evaluation and generative data completion, the dataset annotation quality and model generalization ability are dynamically optimized to solve the problem of model performance degradation caused by data staticization in traditional methods;
[0045] In addition, a path weight allocation strategy combining Bayesian optimization and entropy constraints is adopted to balance the influence of text, structure and context features on the detection results, avoiding false detection caused by the dominance of a single modality; the deep integration of the expert scoring mechanism and the knowledge graph ensures that the model decision-making complies with the professionalism and authority of power safety regulations. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0047] Figure 1 This is a technical flow chart of an embodiment of the present invention for realizing automatic detection of power safety violations;
[0048] Figure 2It is the technical flow chart of step (1) for automatically detecting power safety violation behaviors in the embodiments of the present invention;
[0049] Figure 3 It is the technical flow chart of step (2) for automatically detecting power safety violation behaviors in the embodiments of the present invention;
[0050] Figure 4 It is the technical flow chart of step (3) for automatically detecting power safety violation behaviors in the embodiments of the present invention;
[0051] Figure 5 It is the technical flow chart of step (4) for automatically detecting power safety violation behaviors in the embodiments of the present invention;
[0052] Figure 6 It is the technical flow chart of step (5) for automatically detecting power safety violation behaviors in the embodiments of the present invention;
[0053] Figure 7 It is the technical flow chart of step (6) for automatically detecting power safety violation behaviors in the embodiments of the present invention;
[0054] Figure 8 It is the dataset architecture diagram for automatically detecting power safety violation behaviors in the embodiments of the present invention;
[0055] Figure 9 It is the technical flow chart of step (8) for automatically detecting power safety violation behaviors in the embodiments of the present invention;
[0056] Figure 10 It is the framework schematic diagram of the BERT model in the embodiments of the present invention;
[0057] Figure 11 It is the example diagram of the dataset after positive and negative sample pairing in the embodiments of the present invention. Detailed implementation manners
[0058] To make the features and advantages of the present invention more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0059] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0060] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0061] As Figures 1 - 11 shown, an embodiment of the present invention provides an automatic detection method for power safety violation behaviors, and the specific steps are as follows:
[0062] Step (1): Collect multi-perspective and multi-scene c images covering power operation personnel and their violation behaviors, and use geometric transformation and image scaling techniques to unify all images to a specified size , specifically:
[0063] Target setting: Define the types of power operation scenarios (c) to be collected and the different perspectives to be captured in each scenario ( ), including but not limited to high-altitude operations, substation maintenance, etc.;
[0064] Equipment preparation: Select appropriate photography and videography equipment to ensure that high-quality images can be taken from different angles. Consider using various equipment such as drones, hand-held cameras, or fixed cameras;
[0065] On-site shooting: Dispatch a professional team to different power operation sites to collect images according to the preset perspectives and scenarios. Ensure that as many scenarios and perspectives as possible are covered to obtain a comprehensive data set;
[0066] Simulation environment shooting: In cases where actual operation images cannot be directly obtained, construct a simulation environment for shooting to ensure the diversity of the data set.
[0067] Initial screening and classification: In the initial screening and classification stage, first perform quality detection and evaluation on all the collected images. Use automated scripts to detect the quality of each picture, specifically including blurriness (preferably measured by calculating the Laplacian variance) and exposure (preferably evaluated using histogram analysis). For blurriness, if the Laplacian variance of a picture is lower than the set threshold, the picture is considered too blurry and not suitable as a training sample; for exposure, by analyzing the brightness histogram of the image, judge whether it is within a reasonable range to avoid overexposure or underexposure. In addition, other quality issues need to be checked, such as whether there are obvious noise points or color distortion in the image.
[0068] After completing the quality inspection, classify and organize the remaining high-quality images according to different scenarios (such as high-altitude operations, substation maintenance, etc.) and perspectives (such as top-down, bottom-up, side view, etc.). Ensure that the images under each scenario and perspective are correctly classified, and establish a clear file structure for subsequent annotation and processing work. This process not only helps to improve the overall quality of the dataset but also lays a solid foundation for the training of multi-modal object detection models in the future.
[0069] Construct an image set: Let there be a finite set of perspectives and a finite set of scenarios . Then, all multi-perspective and multi-scenario image sets can be represented in the form of a multi-dimensional array or matrix:
[0070]
[0071] where represents the image under the i-th perspective and the j-th scenario ;
[0072] Size normalization: Use image processing software (such as OpenCV, PIL, etc.) to apply geometric transformations (rotation, flipping, cropping, etc.) and scaling techniques to each image to adjust all images to a unified target size . When performing image scaling, pay attention to keeping the ratio of the original image unchanged to avoid distortion. If necessary, padding can be added to the image edges to reach the target size;
[0073] Data augmentation: To increase the diversity of model training, data augmentation operations can be performed on the adjusted images, such as random cropping, color jittering, noise addition, etc. The processed images are properly saved by category and perspective, and a clear file structure is established for easy access and use in subsequent steps.
[0074] Step (2) Design a series of language descriptions for power violation behaviors , ranging from simple behavior labels to detailed scenario descriptions , ensuring description flexibility, specifically including:
[0075] Clarify the purpose of generating descriptions and the requirements to be met: Short behavior labels and detailed scenario descriptions are needed.
[0076] Professional term extraction: Collect professional terms related to power operations, including types of violation behaviors (such as "not wearing a safety helmet"), locations of occurrence (such as "inside the substation"), equipment involved (such as "high-voltage switch"), risk levels (such as "high risk"), etc.;
[0077] Extract high-frequency words using natural language processing techniques: such as word frequency statistics, TF-IDF analysis, and extract high-frequency words from literature and specification documents in the field of power safety;
[0078] Introduce context-sensitive words: for example, according to different weather conditions (sunny, rainy), time (day, night), or operation status (maintenance, overhaul), dynamically adjust the keywords in the description;
[0079] Design behavior tags : Design short behavior tag templates to directly describe the core actions of violation behaviors, such as "${type} type of violation behavior", "${type} type of violation behavior occurring at ${location}", "${type} type of violation behavior in ${state} state", and replace the placeholders (such as ${type}, ${location}, ${state}) to generate diverse tags; where ${…} is used to represent placeholders, which is a marking method for template variables, used to refer to specific content or values. When creating specific descriptions or sentences, these placeholders will be replaced by actual data or information.
[0080] Expand behavior tags into scenario descriptions : Expand behavior tags and generate detailed descriptions by combining specific scenario information, such as "At ${location}, a worker committed a ${type} type of violation behavior without wearing ${equipmen}, resulting in an increased ${risk} risk.", "${worker} committed a ${type} type of violation behavior when performing tasks at ${location} at ${time}, violating the ${standard} standard."
[0081] Introduce natural language generation (NLG) models: Utilize the powerful context understanding ability of pre-trained models to automatically generate more diverse and complex descriptions; input existing behavior tags and scenario description templates into these models to generate more diverse descriptions.
[0082] Collect and organize existing behavior tags and scenario description templates to form a standardized dataset; use the above dataset to fine-tune the T5 model so that it can better understand and generate descriptions related to electricity violation behaviors; input the given basic description template into the fine-tuned model to generate various modified descriptions. For example, input the template: "At ${location}, a staff member committed a ${type} violation behavior, without wearing ${equipment}, resulting in an increased ${risk} risk."; Output example 1: "In the substation, an electrician committed a violation of not wearing a safety helmet, increasing the risk of head injury."; Output example 2: "In the power distribution room, technicians operated without wearing protective helmets as required, posing a serious safety hazard."
[0083] Evaluate and verify the generated descriptions to ensure their accuracy and applicability. Continuously optimize the model parameters and generation strategies based on the feedback.
[0084] Test, verify, adjust, and optimize the description template: Test the designed template to evaluate its effectiveness and applicability. The flexibility and accuracy of the template can be verified by simulating different scenarios. Adjust and optimize the template according to the test feedback to ensure that the template can meet the requirements of actual applications. Confirm that the template reaches the expected goal and complete the design.
[0085] Step (3) Generate descriptions based on the knowledge graph. By defining key semantic roles, constructing a structured knowledge graph, and extracting path weights, select appropriate description templates to generate diverse descriptions. Specifically:
[0086] Define the key semantic roles in violation behaviors: including the subject (Who), action (What), object (Where / Which), status (How), etc. Specific examples are shown in Table 1 below:
[0087] Table 1 Examples of Key Semantic Roles in Violation Behaviors
[0088] main body staff member electrician technician action not wearing illegal operation entering the prohibited area object safety helmet high - voltage switch prohibited entry area status not in accordance with the regulations unauthorized risk of falling from height
[0089] According to different combinations of semantic roles, dynamically generate language descriptions. For example, the subject is a staff member; the action is not wearing; the object is a safety helmet; the status is not in accordance with the regulations → the description is "The staff member did not wear a safety helmet as required."
[0090] Construct a structured knowledge graph:
[0091] Hierarchical modeling: The knowledge graph is divided into three layers: the entity layer, the event layer, and the impact layer.
[0092] Entity Layer (Entity Layer):
[0093] Personnel: including staff, maintenance personnel, technicians, etc.
[0094] Equipment: such as high-voltage switches, circuit breakers, insulating gloves, protective clothing, etc.
[0095] Location: For example, substations, distribution rooms, power facilities, etc.
[0096] Event Layer:
[0097] Violations: failure to wear a safety helmet, illegal operation of high-voltage switches, failure to wear a seat belt, etc.
[0098] Compliance operations: Wear a safety helmet correctly, operate high-voltage switches according to regulations, use insulating tools correctly, etc.
[0099] Impact Layer:
[0100] Risk: high risk, medium risk, low risk, etc.
[0101] Consequences: risk of electric shock, falling from height, fire hazard, etc.
[0102] Regulatory provisions: violation of safety standards, compliance with safety regulations, exemption clauses, etc.
[0103] Expanded relationship types: Based on the original simple relationships, multiple relationship types are added to more accurately describe complex causal chains and spatiotemporal relationships.
[0104] Causal Relations:
[0105] Directly Leads To: Indicates that one event directly leads to the occurrence of another event.
[0106] Indirectly Triggers: Indicates that one event indirectly triggers another event through a series of intermediate events.
[0107] Temporal and Spatial Relations:
[0108] Occurs At: Indicates the place where the event occurs.
[0109] Lasts Until: Indicates the duration of an event.
[0110] Compliance Relations:
[0111] Violates: Indicates that a certain behavior violates specific safety standards or regulations.
[0112] Complies With: Indicates that a certain behavior complies with specific safety standards or regulations.
[0113] Exempts From: Indicates that in certain situations, compliance with specific safety standards or regulations can be exempted.
[0114] Path weight calculation:
[0115] Node vectorization and similarity calculation: Use a pre-trained word vector model (such as BERT) to generate vector representations for each node. To improve domain adaptability, when fine-tuning the BERT model, add power safety texts (such as the full text of the "Power Safety Work Regulations" and accident reports), as Figure 10 shown.
[0116] Assume that the text descriptions of each node and are and respectively, then their vector representations are and .
[0117]
[0118]
[0119] v i and v j respectively represent the text feature vectors of two nodes (such as entities or events in a knowledge graph), and BERT( ) is used to generate text feature vectors. These vectors are usually generated by a pre-trained word vector model (such as BERT) and are fine-tuned for domain adaptability to enhance the understanding ability in the power safety field.
[0120] Utilize the structural information of nodes in the knowledge graph to enhance similarity calculation. Specifically, the structural features of nodes can be extracted through a graph neural network (GNN). Assume that and are and respectively.
[0121]
[0122]
[0123] s i and s jRepresents the structural feature vectors of two nodes. GNN( ) is used to extract the structural features of nodes from the knowledge graph. This is usually extracted from the knowledge graph through a graph neural network (GNN), which reflects the position of the node in the network and its relationship with other nodes.
[0124] Taking into account the context information of the nodes (such as time, location, etc.), the similarity calculation can be further enhanced. Suppose and are the context information of nodes and respectively, then their vector representations are and .
[0125]
[0126]
[0127] c i and c j refer to the context feature vectors, which include information such as time, location, and operation status, and are generated by a spatio-temporal encoder. ContextEncoder( ) is responsible for encoding context information such as time, location, and operation status into vector representations.
[0128] Combining the above three features (text, structure, context), the comprehensive similarity is calculated. The specific formula is as follows:
[0129]
[0130] Where: Represents the cosine similarity between two vectors; Are the weight coefficients of the text feature, structural feature, and context feature respectively, satisfying . Specifically, according to the knowledge of domain experts, the importance of each feature is initially evaluated and the initial weights are set. Using the Bayesian optimization method, the weight coefficients are dynamically adjusted during the training process. Design a feedback-based learning and correction mechanism to update the model parameters in real time to adapt to new data. Construct a joint loss function to optimize the interaction between multiple features simultaneously. Since the weights may exceed the reasonable range due to no constraints during the Bayesian optimization process (such as →1, completely ignoring the structural and context features), each weight coefficient is forced to be ≥ 0.2 to ensure that the multi-modal features participate evenly. And add a weight entropy penalty term to the objective function of Bayesian optimization to encourage a uniform weight distribution: where represents the regularization loss function; λ is a hyperparameter called the regularization strength coefficient; k ∈ {t, s, c}, where k is an index variable representing an element in the set {t, s, c}, and t, s, and c represent the weights of temporal features, structural features in the knowledge graph, and context features respectively; for each k, w k represents the corresponding weight value. These weights usually come from the parameters learned during the model training process, and they reflect the importance or contribution degree of different types of features.
[0131] Weight formula optimization: The weights can be calculated by the following formula, which combines node similarity, relationship frequency, expert scores, and semantic relevance between nodes:
[0132]
[0133] where is the similarity between nodes and ; is the frequency of relationship , that is, traversing the entire knowledge graph to count the number of times a specific relationship appears as an edge in the graph; is the semantic relevance score between two nodes, used to measure their semantic association degree; the parameter β is a parameter used to adjust the steepness of the curve; is the normalized expert score that maps the expert score to the interval [0, 1]; and are weight adjustment functions (such as linear or non-linear functions). The f function is used to adjust the similarity score, and it is preferably to use the Sigmoid function to amplify the difference, making the high similarity score more prominent. The g function is used to adjust the weight according to the relationship frequency, and it is preferably to adopt logarithmic normalization processing; h is a function of semantic relevance used to measure the semantic association degree between nodes.
[0134]
[0135] where expert_score is the expert score, and min_score and max_score correspond to the minimum and maximum values in the expert score respectively.
[0136]
[0137] is the input similarity score; α is a parameter that controls the steepness of the curve. A larger α value will make the function change more violently, thus being more sensitive to the change of the similarity score.
[0138]
[0139] y is the relation frequency of the input; is the maximum frequency value among all the relations considered.
[0140]
[0141] in Is a node and The semantic relevance score between is a parameter that controls the steepness of the curve. The semantic relevance score can be obtained by calculating the semantic distance between two node descriptions using a pre-trained language model (such as BERT).
[0142] Template matching and generation:
[0143] Template design: define templates, such as Template 1: In ${location}, ${who} caused ${consequence} (violating ${standard}) due to ${action}${object}; Template 2: ${worker} did not follow the prescribed ${action} when performing an operation in ${location} at ${time}, increasing the risk of ${consequence};
[0144] Synonym replacement: increased risk → trigger hidden dangers, increase threats; not worn → not worn, not worn correctly, not worn according to regulations;
[0145] Multi-path fusion:
[0146] Attention mechanism: Calculate the normalized weight of each path Where 𝑊 is the sum of all path weights. Weighted average description of each path.
[0147] Conflict resolution rule: Select the consequence corresponding to the path with the highest weight as the final result.
[0148] Post-processing verification:
[0149] Fluency correction: Use a language model (such as GPT-2) to correct the fluency of the generated text.
[0150] Dependency parsing: ensuring the subject, predicate and object structure is correct.
[0151] Step (4) semantic reorganization and detail supplementation: by replacing key terms in the description with synonyms, adjusting sentence structure, reorganizing the order of information, and adding more contextual information to make the description more specific and detailed. Specifically:
[0152] Synonym replacement: Use a term mapping table or a pre-trained language model to perform synonym replacement and generate new description variants. The specific steps are as follows:
[0153] Construct a term mapping table dedicated to the power field, including but not limited to the content in Table 2:
[0154] Table 2 Special Terms in the Power Field and Their Synonyms / Related Expressions
[0155] terminology synonyms / related expressions safety helmet helmet, protective helmet, work helmet insulating gloves protective gloves, work gloves, insulating equipment protective clothing work clothes, insulating clothing, protective equipment substation distribution room, power facilities, high - voltage station, power supply station electrician maintenance worker, technician, staff member, maintenance personnel daily maintenance work equipment inspection, routine maintenance, daily patrol inspection, maintenance operation not wearing not wearing, not wearing correctly, not wearing in accordance with the regulations not fastened not fastened correctly, not fixed, not tied tightly working at height working at elevation, high - altitude construction, working at height risk of electric shock electrical hazard, current injury, risk of electric shock high risk extremely high risk, major hidden danger, serious threat medium risk relatively high risk, general hidden danger, medium threat low risk minor hidden danger, low - level threat, controllable risk illegal operation improper operation, wrong operation, violation of regulations insulating tools protective tools, safety appliances, insulating equipment emergency passage evacuation route, escape route, emergency exit prohibited entry area hazardous area, restricted area, isolation area high - voltage switch circuit breaker, distribution switch, high - voltage equipment personal protective equipment (PPE) protective articles, safety equipment, labor protection articles falling from height falling accident, falling from height, risk of falling risk of fire fire hazard, fire threat, risk of thermal runaway unauthorized unqualified, uncertified, unpermitted maintenance record maintenance log, operation record, equipment ledger emergency plan emergency handling plan for unexpected events, emergency response plan, accident response measures
[0156] For the key terms in the input description, replace them with synonyms one by one. For example, for the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One electrician is not wearing a safety helmet, and another electrician is not wearing a seat belt." Replacement rules: safety helmet → helmet, substation → distribution room, electrician → maintenance worker;
[0157] The generated description after replacement: Let the input description be D, the term set be , and the synonym set be , s ik represents the kth synonym of the ith term, then the description after replacement can be expressed as: , where Replace is the replacement function, where D is the description, t i is the ith term.
[0158] For example, for the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One electrician is not wearing a safety helmet, and another electrician is not wearing a seat belt.", Replacement result 1: "In a certain distribution room, multiple maintenance workers are performing daily maintenance work. One maintenance worker is not wearing a helmet, and another maintenance worker has not properly fastened the protective equipment."; Replacement result 2: "In a certain power facility, multiple technicians are performing equipment inspections. One technician is not wearing a protective helmet, and another technician is not wearing work clothes as required."
[0159] Context awareness: Extract context information from the input description, such as weather conditions, time, operation status, etc.; Dynamically adjust the keywords and expressions in the description according to the extracted context information.
[0160] For example, for the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One electrician is not wearing a safety helmet, and another electrician is not wearing a seat belt."; The dynamically adjusted description: "On a sunny day, in the substation, multiple electricians are performing daily maintenance work. Due to not wearing a safety helmet and not wearing a seat belt, one electrician may face the risk of head injury, and another electrician has the potential hazard of falling from a height."
[0161] Semantic Reorganization: By adjusting the sentence structure or reorganizing the information order, generate descriptions with the same semantics but different expressions.
[0162] Adjusting the sentence structure or reorganizing the information order: Subject-predicate-object adjustment, that is, changing the order of the subject, predicate, and object in the sentence; Clause nesting, that is, nesting some information into clauses; Parallelism and progression, that is, using parallel sentences or progressive sentences to enhance the logic of the description.
[0163] Let the input description be D, the sentence structure be , and S be a set containing three elements: Subject, Verb, and Object. Then the reorganized description can be expressed as: , where Reorganize is the semantic reorganization function.
[0164] For example, the input description: "In a certain substation, multiple electricians are performing daily maintenance work. One of the electricians is not wearing a safety helmet, and another electrician is not wearing a seat belt." Reorganization result 1: "Due to not wearing a safety helmet and not wearing a seat belt, there is a risk of falling from a height for multiple electricians working in a certain substation." Reorganization result 2: "In a certain substation, although multiple electricians are performing daily maintenance work, one of the electricians is not wearing a helmet correctly, and another electrician is not wearing the seat belt properly as required."
[0165] Adding Detail Supplements: By introducing more context information or expanding the description content, make the generated description more specific and detailed.
[0166] Introducing more context information or expanding the description content: Supplementing the risk level, that is, clarifying the specific risks that illegal acts may cause; Time and location, that is, adding the time and specific location of the event; Equipment information, that is, supplementing the names or models of the equipment involved.
[0167] Let the input description be D, and the supplementary information be , and C be a set containing four elements: Risk, Time, Location, and Equipment. Then the supplemented description can be expressed as: , where Expand is the supplement function.
[0168] For example, if you input the description: "In a substation, several electricians are performing routine maintenance work. One of them is not wearing a safety helmet, and another is not wearing a safety belt.", Supplementary result 1: "In a substation (number #123), several electricians are performing routine maintenance work on high-voltage lines at 9 a.m. One of them is not wearing a protective helmet, which may cause head injury; another electrician does not wear a safety belt as required, which poses a risk of falling from a height.", Supplementary result 2: "In a substation, several electricians are performing maintenance work on transformer equipment. One of them is not wearing an insulating helmet correctly, which may cause an electric shock accident; another electrician does not wear protective equipment as required, which increases the safety hazard of the overall operation."
[0169] Generate composite descriptions by combining Boolean logic: Generate composite descriptions containing multiple violations based on the combination relationship of multiple violations.
[0170] Let the set of illegal behaviors be , the logical relationship is L, then the composite description can be expressed as: , where Combine is the combining function.
[0171] For example, if we input the violation behavior set: B = {not wearing a helmet, not wearing a safety belt, not wearing protective clothing}, the logical relationship is: , output description: "In a substation, several electricians are doing routine maintenance work. One of the electricians neither wears a safety helmet nor fastens his safety belt correctly. At the same time, another electrician does not wear protective clothing as required, resulting in multiple safety hazards in the overall operation."
[0172] Step (5) annotates the traffic violation in each image in detail, including bounding box positioning, instance segmentation mask, and behavior state labeling, and combines text-based image generation technology to supplement missing data. Specifically:
[0173] Bounding box positioning: Use professional image annotation tools (such as LabelImg or CVAT) to draw an accurate bounding box for each violation instance and record its coordinate position ( , , , ), where x min ,y min Respectively represent the horizontal and vertical coordinates of the lower left corner of the rectangle, x max ,y max Respectively represent the horizontal and vertical coordinates of the upper right corner of the rectangle. At the same time, the bounding box is combined with the behavior label generated in step S2. associations, such as “not wearing a safety helmet” or “illegal operation of a high-voltage switch”;
[0174] Instance Segmentation Mask: Generate pixel-level segmentation masks using manual segmentation tools (such as LabelMe) or in combination with pre-trained instance segmentation models (such as Mask R-CNN) to ensure that every detail of the target area is covered. Combine the mask with a more detailed scene description corresponding to, for example, "Inside the substation, the staff did not wear protective equipment as required and carried out dangerous operations";
[0175] Add detailed status information to each instance of violation behavior: including but not limited to timestamp, location information, number of persons involved, etc.
[0176] Language-Visual Association: Ensure the consistency and accuracy of annotations by quantifying the correlation between images and texts. For example, use a pre-trained CLIP model to extract image and text features to calculate the similarity. If the similarity is lower than a threshold (such as 0.75), determine that the image and text do not match and trigger the re-annotation process;
[0177] Check for missing pictures: After completing the above annotation and language-visual association operations, if it is found that some descriptions do not have corresponding pictures, corresponding pictures can be generated through text-to-image synthesis technology.
[0178] Preprocess the input descriptions (descriptions without corresponding pictures): Use a BERT-based NER model to identify entities (locations, persons, actions, objects) in the descriptions; combine with an external knowledge base (power safety documents) to dynamically add details: such as "The staff did not wear a safety helmet" → enhanced to: "An electrician in work clothes did not wear a safety helmet when conducting equipment inspections inside the substation on a sunny day."
[0179] Generate images based on text: Input the enhanced descriptions into Stable Diffusion to generate images: Select a resolution of 1024×1024 (high detail is required for power scenarios); realistic style (disable artistic style); generate 3 - 5 candidate images for screening.
[0180] Further optimize the generated images: Use OpenCV to adjust the brightness / contrast to ensure consistency with the distribution of real images, and screen the image with the highest text similarity based on CLIP.
[0181] Quality Control: Calculate the FID score (Frechet Inception Distance), requiring that the FID between the generated image and the real image ≤ 15; verify the semantic accuracy of the generated image (such as the safety helmet model, scene compliance) by power safety experts.
[0182] Step (6) Create positive and negative sample pairs based on language descriptions , where the positive samples correspond to accurately described behaviors, while the negative samples contain behaviors that do not match or are irrelevant to the description. Specifically:
[0183] Select a matching description: According to the type and detailed information of the violation behavior marked in the image, select the most appropriate language description as the positive sample description. For example, if the image shows an electrician not wearing a safety helmet, select a description such as "Inside the substation, an electrician is performing maintenance work without wearing a safety helmet as required";
[0184] Generate positive sample pairs: Pair the image with its corresponding correct description to form positive sample pairs (image, description), where image refers to the image data containing a specific violation behavior; description is the text description that exactly matches the above image and accurately reflects the violation behavior in the image. Ensure that each pair matches exactly to avoid any inconsistent situations;
[0185] Generate non-matching descriptions: Generate non-matching descriptions by means of random replacement, deliberate errors, and scene dislocation: Randomly select descriptions from the language description library that are irrelevant to the violation behavior of the current image, or deliberately select descriptions that are opposite to the type of violation behavior in the image (such as if the image shows "not wearing a safety helmet", the description is "wearing a safety helmet correctly"), and select descriptions of another scene (such as if the image shows working at height, the description is "operating inside the substation");
[0186] Introduce the generation of adversarial samples:
[0187] Image adversarial samples: Use adversarial attack algorithms (such as FGSM, PGD, etc.) to add tiny perturbations to the image, making the image look almost the same, but causing the model to produce misclassifications or descriptions. Use adversarial attack algorithms to generate images with tiny perturbations and assign the same description as the original image to them. Such samples can test the sensitivity of the model to subtle changes.
[0188] Description adversarial samples: Make the description seem semantically reasonable but actually not match the image by changing some keywords or phrases in the description. For example, keep most of the description unchanged but make misleading modifications in the key parts. Make subtle semantic adjustments to the original description to generate a description that seems reasonable but is actually inconsistent. For example, if the image shows "not wearing a safety helmet", the description is "Although not wearing a safety helmet, other protective equipment is complete".
[0189] Generate negative sample pairs: Pair an image with a mismatched description to form negative sample pairs (image, incorrect_description). Ensure that each pair is mismatched to train the model to distinguish correct descriptions from incorrect ones, where image refers to the image data containing specific violation behaviors; incorrect_description is a deliberately mismatched description, i.e., the content it describes does not match the actual situation in the image.
[0190] Image adversarial sample pairs: Pair an image with a small perturbation with its original description to form adversarial sample pairs (adversarial_image, original_description), where adversarial_image is an image generated by adding a small perturbation to the original image through an adversarial attack algorithm (such as FGSM, PGD, etc.); original_description is the original correct description corresponding to the adversarial sample image, and even if the image has been perturbed, its description should remain unchanged.
[0191] Description adversarial sample pairs: Pair the original image with its description after semantic adversarial adjustment to form adversarial sample pairs (original_image, adversarial_description), where original_image refers to the real-scene image without any modification or perturbation, which accurately reflects a specific power operation scene and its possible violation behaviors; adversarial_description is a description generated through specific processing to make it inconsistent with the content of the original image or contain misleading information.
[0192] To ensure the quality of positive and negative sample pairs, strict review and verification are carried out: including cross-checking, expert review, and establishing a feedback mechanism to continuously improve the accuracy and consistency of the sample pairs.
[0193] Step (7) Integrate the results of the previous steps to construct a dataset of power safety violation behaviors that comprehensively covers various scenarios and complex language descriptions , specifically:
[0194] Summarize and organize all the data generated in the previous steps, including images, annotation information, language descriptions, and their corresponding positive and negative sample pairs:
[0195] Image data: All annotated power violation images, including bounding box localization, instance segmentation masks, and behavior status markers;
[0196] Language descriptions: Cover various language description templates from short behavior labels to detailed scene descriptions;
[0197] Positive and negative sample pairs: Make sure that each image has a corresponding positive sample pair (accurately described behavior) and negative sample pair (inconsistent or irrelevant behavior), such as Figure 11 As shown, the positive sample description P given by the corresponding picture I There are: "Workers wearing blue helmets", "Workplaces with safety fences", "Workers next to a white truck are on the phone"; negative sample description N I There are: "There are no workers wearing safety helmets", "There are no illegal operations in the picture";
[0198] Step (8) trains a multimodal detection model by quantifying the correlation between images and texts and continuously optimizes the dataset labels to improve the accuracy and robustness of the model. Specifically:
[0199] Dataset division: divide the dataset into training set, validation set and test set to ensure balanced distribution;
[0200] Model selection: Select ViLBERT, a model architecture suitable for multimodal tasks, for training. This model can process image and text inputs simultaneously and learn the correlation between them.
[0201] Loss function design: Design a suitable loss function (such as contrast loss, cross entropy loss, etc.) to guide model training. The following is the design of the specific loss function:
[0202] Contrastive Loss: The goal of contrastive loss is to bring positive pairs closer together while pushing negative pairs further apart. Assumptions: The feature vector representing the image x; The feature vector representing the text description y; represents the Euclidean distance between the two; m>0 is a predefined threshold used to control the distance between negative sample pairs. The contrast loss formula is:
[0203] Where N is the total number of samples; is a label indicating whether the sample pair is a positive sample pair ( ) or negative sample pairs ( );
[0204] Cross-Entropy Loss: Cross-Entropy Loss is used for supervised classification tasks to calculate the difference between the model's predicted distribution and the true distribution. Assumptions: is the class probability distribution predicted by the model; is the one-hot encoding distribution of the true label. The cross entropy loss formula is:
[0205] where N is the total number of samples; C is the total number of categories; is the sample true label; is the probability distribution predicted by the model.
[0206] Combined loss function: Combine the contrastive loss and the cross-entropy loss to form a combined loss function. Assume: and are two hyperparameters used to balance the importance of the two losses. The formula for the combined loss function is: . If more attention is paid to the matching relationship between the image and the text, then increase ; if more attention is paid to the accuracy of the classification task, then increase .
[0207] Feature extraction: Use the visual encoder of ViLBERT to extract feature vectors from images; use the language encoder of ViLBERT to extract feature vectors from text descriptions.
[0208] Calculate the relevance between the image and the text to evaluate their degree of correlation: Use cosine similarity, dot product similarity or other similarity measurement methods to quantify the relevance between the image and text features. For example, for each image-text pair, calculate the cosine similarity score between their feature vectors.
[0209] Optimize model parameters: Through the backpropagation algorithm, adjust the model parameters according to the gradient of the contrastive loss function, so that the similarity score between positive sample pairs is maximized and the similarity score between negative sample pairs is minimized.
[0210] Regularly evaluate model performance: Regularly evaluate model performance (such as accuracy, recall rate, F1 score), and adjust hyperparameters, increase training data or improve the model architecture according to the evaluation results. For example, if the model performs poorly on certain specific types of violation behaviors, the model effect can be improved by increasing the data of these types.
[0211] Model deployment: Deploy the trained multi-modal model to the actual application scenario for real-time identification of power violation behaviors. This model can make full use of the design of the previous series of high-quality data sets to achieve accurate identification and efficient management of violation behaviors in power operations.
[0212] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.
[0213] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0214] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0215] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
[0216] The present invention is not limited to the above best implementation manners. Anyone can obtain various other forms of automatic detection methods for power safety violation behaviors under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.
Claims
1. An automatic detection method for power safety violation behaviors, characterized in that: Construct a multi-modal dataset covering power operation scenarios through multi-view image acquisition and knowledge-graph-driven dynamic semantic recombination, and the dynamic semantic recombination generates diverse language descriptions based on the hierarchical structure and path weight optimization of the knowledge graph; Expand the dataset based on adversarial sample generation technology to generate image perturbation samples and semantic misleading text descriptions; Train an object detection network through a multi-modal feature dynamic fusion model, and the model realizes image-text correlation learning through knowledge graph path weight optimization and joint loss function; The comprehensive similarity calculation formula for the path weight is: Wherein: represents the cosine similarity between two vectors; are the weight coefficients of text features, structural features, and context features respectively, satisfying ; the weights are dynamically optimized by Bayesian search; v i and v j respectively represent the text feature vectors of two nodes; s i and s j respectively represent the structural feature vectors of two nodes; c i and c j refer to the context feature vectors, which are generated by the spatio-temporal encoder; The calculation of the path weight fuses the relationship frequency factor and expert scores, and the formula is: where is the similarity between nodes and ; is the frequency of the relationship ; is the semantic relevance score between two nodes; the parameter β is a parameter used to adjust the steepness of the curve; is the normalized expert score mapping the expert score to the interval [0, 1]; and are weight adjustment functions. The f function is used to adjust the similarity score, and the g function is used to adjust the weight according to the relationship frequency; h is a function of semantic relevance, which is used to measure the semantic association degree between nodes; Infer the real-time collected power operation images based on the trained model, and output the detection results of violation behaviors and their semantic descriptions; Trigger a closed-loop feedback through multi-modal similarity evaluation to optimize dataset annotation and model parameters.
2. The automatic detection method for power safety violation behaviors according to claim 1, characterized in that: The hierarchical structure of the knowledge graph includes: The entity layer, including entity of personnel, equipment and location; The event layer, connecting violation behaviors and risk consequences through direct causation, occurrence in and violation relationship types; The impact layer, defining risk levels, associated regulatory clauses and consequence descriptions.
3. The automatic detection method for power safety violation behaviors according to claim 1, characterized in that: The dynamic semantic recombination includes: Synonym replacement: Generate diverse description variants based on the power safety term mapping table, and enhance semantic coherence through a natural language generation model; Boolean logic generation: Based on the logical relationships in the event layer of the knowledge graph, generate compound risk scenario descriptions by combining multiple violation behaviors; Situation awareness: Dynamically expand time, weather and equipment information to generate detailed descriptions; Scene description generation: Dynamically combine violation behaviors, risk consequences and context information based on the causal relationship chain and spatio-temporal relationship chain of the knowledge graph.
4. The automatic detection method for power safety violation behaviors according to claim 3, characterized in that: The natural language generation model is a T5 model fine-tuned based on power safety domain texts, generating context-adapted synonym combinations; The compound risk scenario description combines violation behaviors through logical operators AND / OR, and associates with the risk consequences in the impact layer of the knowledge graph; The situation awareness dynamically extracts time, location and equipment information through the spatio-temporal relationship chain of the knowledge graph to generate detailed descriptions.
5. The automatic detection method for power safety violation behaviors according to claim 1, characterized in that: The knowledge graph path weight optimization includes: Node vectorization: Generate text feature vectors based on BERT fine-tuned through power safety specification texts, and generate structural feature vectors through GNN; Context feature extraction: Generate time, location and operation status feature vectors through a spatio-temporal encoder.
6. The automatic detection method for power safety violation behaviors according to claim 1, characterized in that: The adversarial sample generation technology includes: Add perturbation momentum to the original image through the FGSM or PGD algorithm to generate visual adversarial samples; Generate text adversarial samples by replacing keywords or adjusting logical relationships.
7. The automatic detection method for power safety violation behaviors according to claim 1, wherein: The closed-loop feedback includes: Use the CLIP model to calculate the image-text similarity score, and trigger the re-annotation process when it is lower than the threshold; Generate missing scene images through Stable Diffusion, constrain the FID score of the generated images ≤ 15, and verify compliance through power safety experts.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1-7.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Operation site safety management and control method and system based on multi-modal large model
CN119625477A