A substation violation detection method based on real-time open vocabulary detection

By improving the YOLO architecture and using a real-time open vocabulary detection method based on multimodal feature fusion, the problems of untimely response and insufficient adaptability in substation violation detection are solved, achieving efficient and accurate identification and localization of violations, which is suitable for edge computing environments.

CN120510571BActive Publication Date: 2025-11-04XIAMEN UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511008626.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-04
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

Existing substation violation detection technologies rely on manual judgment, resulting in untimely responses and limited coverage. Furthermore, traditional deep learning models struggle to adapt to the identification of complex and diverse violations, and the requirements for lightweight and real-time performance in edge computing environments are difficult to meet.

Method used

By adopting a real-time open vocabulary detection method, and by improving the YOLO architecture and multimodal feature fusion, combined with semantic abstraction and feature compression, we can achieve efficient identification and localization of violations in substations and perform real-time detection on edge devices.

Benefits of technology

It improves the accuracy and robustness of detection, adapts to complex power operation environments, has cross-modal feature fusion capabilities, supports seamless detection of open categories, and enhances the real-time performance and engineering scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510571B_ABST
    Figure CN120510571B_ABST
Patent Text Reader

Abstract

The application discloses a substation violation detection method based on real-time open vocabulary detection, comprising the following steps: collecting substation field operation images, combining operation specification documents to extract the violation description prompt sentences to be detected, and coding the violation description prompt sentences into prompt sentence vectors; inputting the field operation images into a feature extraction network of an improved YOLO architecture to obtain multi-scale image semantic features; for the detection task, performing decoupling text adaptive processing on the semantic vectors of the violation description prompt sentences to obtain enhanced semantic vectors; fusing and aligning the multi-scale semantic features and the enhanced semantic vectors to obtain multi-modal features, and performing open category violation behavior recognition and positioning through the multi-modal features; embedding and caching similar and repeated violation description prompt sentences, and deploying the same on an edge device to perform real-time violation behavior detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of substation violation detection, and particularly relates to a substation violation detection method based on real-time open vocabulary detection. BACKGROUND

[0002] With the continuous improvement of the intelligence level of the power system, the substation plays an increasingly important role in the safe operation of the power grid. In order to protect the stability and reliability of the power system, real-time and accurate safety supervision of the operation behavior of the substation has become a key task of the power industry. However, most of the current substations still mainly rely on manual patrol or traditional video monitoring based on rules, and the identification of violation behaviors depends on manual judgment, which has many problems such as untimely response, limited coverage, high misjudgment and omission rate, and is difficult to meet the intelligent detection needs of high efficiency and high accuracy in modern power operation scenarios. In recent years, with the rapid development of deep learning technology, target detection algorithms based on convolutional neural networks (CNN) such as YOLO series and Faster R-CNN have been widely used in various video analysis tasks. Although these methods have achieved good results in fixed-class scenarios, they lack the ability to understand and generalize new classes or fine-grained semantics due to their dependence on the pre-defined closed set of labels in the training phase, making it difficult to adapt to the complex and diverse violation behavior recognition tasks in substations.

[0003] In addition, the target size in the substation monitoring video changes dramatically, the occlusion situation is serious, and the violation behavior has strong semantic description dependence. For example, "not wearing a reflective vest", "operating in a live area", "private cable pulling", and other complex categories cannot be accurately represented by traditional classification methods, which seriously restricts the application range of existing detection technologies. At the same time, there are higher requirements for model lightweight and real-time in edge computing environment, and traditional deep models are difficult to directly deploy on resource-constrained embedded devices for efficient inference.

[0004] In recent years, the rise of large-scale pre-trained multi-modal models has provided a new way to solve the above problems. This type of method introduces natural language as a target description means, and is no longer limited to a fixed class set, with stronger generalization ability, especially suitable for identifying new classes that do not appear in the training set. However, the current mainstream open vocabulary detection methods still have problems such as inaccurate semantic matching, text ambiguity interference, low efficiency of image-text fusion, and difficulty in edge deployment, which limit their practical application effect in industrial scenarios.

[0005] Aiming at the problems existing in the prior art, a substation violation detection method based on real-time open vocabulary detection is designed, which is the research purpose of the present application. SUMMARY

[0006] Therefore, the present application aims to provide a substation violation detection method based on real-time open vocabulary detection, which can solve the above problems.

[0007] The present application provides a substation violation detection method based on real-time open vocabulary detection, comprising:

[0008] Collecting substation field operation images, combining operation specification documents to extract violation description prompt sentences for detection, and encoding the violation description prompt sentences into prompt sentence vectors;

[0009] Input the field operation images into the feature extraction network of the improved YOLO architecture to obtain multi-scale image semantic features;

[0010] For the detection task, the semantic vectors of the violation description prompt sentences are subjected to decoupling text adaptation processing to obtain enhanced semantic vectors;

[0011] The multi-scale semantic features and the enhanced semantic vectors are fused and aligned to obtain multi-modal features, and the multi-modal features are used for recognition and positioning of open-class violation behaviors;

[0012] Similar and repeated violation description prompt sentences are embedded and cached, and deployed on edge devices for real-time violation behavior detection.

[0013] The present application has the following advantages:

[0014] First, by extracting violation descriptions, automatic completion of the object-verb structure extraction, dictionary alignment and consistency verification with the operation object label, semantic redundancy, ambiguity and repetition are effectively removed, ensuring that the detection categories strictly correspond to the actual business, improving the detection specificity and accuracy. Form a pair of "violation description-relevant label" input. Through consistency verification, a semantic input pair specific to the task scene is generated, which can adapt to the customized and dynamically expanded detection needs.

[0015] Second, by improving the YOLO network backbone and path enhancement, multi-scale fusion module, multi-scale and cross-level feature expression and aggregation are realized, which has strong perception and modeling ability for small targets, local features and complex scenes, greatly improving the accuracy and robustness of violation detection in complex power operation environment. The bottom-up and top-down information flow promotes the collaborative work of different feature layers, greatly enhancing the scene adaptability and feature representation richness. All scale features are mapped to a shared space, which is conducive to efficient cross-modal alignment with subsequent text features.

[0016] Thirdly, through semantic abstraction, feature compression and nonlinear transformation, and task-guided weighting mechanism, the deep structuring of text semantic information and the task-specific adaptation of detection tasks are realized, and the generalization and fault tolerance for synonymous descriptions, context changes and expression diversity in different detection tasks are higher. It is beneficial to generate enhanced semantic vectors with higher abstraction levels and high matching with target detection tasks, laying a solid foundation for cross-modal feature fusion and open class detection.

[0017] Fourthly, through the fusion gate mechanism and multi-head cross attention, efficient alignment of multi-scale image features and enhanced text semantics is realized, and new descriptions and self-defined texts in open class scenarios can be effectively adapted and recognized. Through feature enhancement, splicing and multi-source information fusion, the overall perception ability of the detection model to spatial distribution, semantic details and target position is strengthened. The spatial distribution shape label is used to distinguish the bounding box, effectively improving the robustness and precision of position judgment and semantic classification under multi-target and multi-class conditions, and supporting seamless detection of newly added classes of text.

[0018] Fifthly, by constructing an embedding vector cache and hash matching mechanism, repeated coding of high-frequency or repeated illegal description sentences is avoided, the high-speed reuse of text semantic features is realized, the detection inference response time is significantly shortened, and the real-time performance of the overall system is improved. Combined with LRU and other cache eviction strategies, the cache space and memory overhead are effectively controlled, and the deployment adaptability and engineering scalability of edge devices are strengthened. Inference acceleration does not affect the core recognition effect, making the detection system lightweight and efficient, suitable for actual industrial site and low-power terminal deployment requirements. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0020] Figure 1 is the method flowchart of the present embodiment.

[0021] Figure 2 is the path enhancement method flowchart of the present embodiment.

[0022] Figure 3 is the image-text joint coding and interaction flowchart of the present embodiment. DETAILED DESCRIPTION

[0023] To facilitate understanding by those skilled in the art, the structure of the present invention will now be described in further detail with reference to the accompanying drawings. It should be understood that, unless otherwise specified, the order of the steps mentioned in this embodiment can be adjusted according to actual needs, and they can even be executed simultaneously or partially simultaneously.

[0024] like Figure 1 As shown, this embodiment of the invention provides a substation violation detection method based on real-time open vocabulary detection, including:

[0025] S1 collects images of on-site operations at the substation, extracts violation descriptions and prompts from the operation specifications document, and encodes the violation descriptions and prompts into a prompt vector.

[0026] S101 acquires images of on-site operations at the substation, extracts violation description prompts from the operation specification documents corresponding to the operation task type, marks the corresponding object tags, and generates task-related semantic input pairs.

[0027] S1011 captures images of substation operations via video recording. ,in, and These represent the height and width of the image, respectively, and 3 indicates a three-channel image;

[0028] In this step, the content of the substation on-site operation images includes: workers, equipment, tools, background areas, etc., and has unstructured characteristics such as high density, multiple targets, and severe occlusion.

[0029] S1012 extracts the description and prompt statements of the violations to be detected from the work specification document corresponding to the work task type. ,in, A set of statements;

[0030] In this step, based on the job task type and operating procedures, at least one violation description prompt statement (natural language prompt statement) is extracted from the standard job description text. This statement is used to indicate the type of violation or object to be detected, such as "operator not wearing a safety helmet" or "personnel entering a live area".

[0031] S1013 processes the violation description prompt statement into a verb-object structure and aligns its semantics with the violation behavior dictionary;

[0032] In this step, the natural language prompt must be a short text with a verb-object structure (such as "detecting no helmet being worn"), and its semantics must match the preset dictionary of violations. Alignment.

[0033] S1014 defines an object label set based on the actual task scenario. wherein each label represents a category word and / or a specific object related to the scene;

[0034] In this step, the object label set is defined according to the actual task scene wherein each label represents a category word or a specific object (such as "safety helmet", "climbing board", "cable", "worker action") related to the scene, and the label information is used to supplement or constrain the detection range. The label set can be derived from a manually predefined category or from a language model pre-training category space . The object label set is used for redundancy elimination and alignment, and also prepares for the category diversity in subsequent multi-modal detection.

[0035] S1015 consistency verification is performed on the violation description prompt sentence and the label set to eliminate semantic repetition and / or category redundancy, and a task-related semantic input pair is generated.

[0036] In this step, consistency verification is performed on the violation description prompt sentence and the label set to eliminate semantic repetition or category redundancy, and a task-related semantic input pair is generated as the guided input of the subsequent image-text fusion module to ensure that the detection target meets both the semantic description and the supervision requirement of the constraint label.

[0037] S102 semantic encoding is performed on each violation description prompt sentence to obtain a corresponding semantic vector, and the object label set is mapped and encoded into a label embedding matrix. The semantic vector and the label embedding matrix are jointly formed into a task-related label subset.

[0038] S1021 the violation description prompt sentence set is encoded by a pre-trained multi-modal language model to generate a semantic vector wherein represents a d-dimensional real number vector space;

[0039] In this step, the violation description prompt sentence set (the set of violation description prompt sentences ) is encoded by a pre-trained multi-modal language model (CLIP or BLIP) to generate a semantic vector wherein the embedding dimension d is unified with the image dimension and is set to 512, and satisfies the normalization condition .

[0040] S1022 If the violation description prompt sentence set is a compound sentence, it is semantically segmented and split into atomic prompts The atomic prompts are encoded as and fused into a unified representation through an attention mechanism.

[0041] In this step, if the violation description prompt sentence set is a compound sentence (such as "detecting whether a safety helmet or insulating gloves is not worn"), it is first segmented and split into atomic prompts , and then encoded as , and finally fused into a unified representation through an attention mechanism.

[0042] S1023 Map each label item in the object label set to the corresponding label vector to form a label embedding matrix , where represents a set of real matrices of n rows and d columns.

[0043] This step is used to assist in describing the category prior and target restrictions of the detection task.

[0044] S1024 Jointly model the semantic vector and the label embedding matrix , calculate the semantic similarity , reorder or filter the label set, and form a task-related label subset .

[0045] This step is used to improve the accuracy and category guidance of the input semantics.

[0046] S1025 If the violation description prompt sentence set is empty, use the category set constructed in the training phase as the target category set.

[0047] In this step, if the violation description prompt sentence set is empty, the detection method is based on the detection mode of the predefined category set, skips the natural language embedding process in the semantic encoding steps S1021-S1024, and directly uses the category set constructed in the training phase as the target category set to participate in subsequent detection. If no detection description sentence is input, the system automatically switches to using the "built-in categories during training" for detection, and no longer uses natural language guidance and open category scheme in the future.

[0048] S2 inputs the field operation image into the feature extraction network of the improved YOLO architecture to obtain multi-scale image semantic features;

[0049] S201 constructs the improved YOLO architecture by the backbone extraction network, the path enhancement module and the multi-scale feature fusion module, and inputs the field operation image into the improved YOLO architecture.

[0050] In this step, the backbone network adopts the feature extraction network of YOLOv8, which is used for multi-scale feature extraction of the input image to generate visual embedding representations of different scales, so as to enhance the detection capability of the target.

[0051] S202 performs multi-level convolution and spatial down-sampling on the field operation image by the backbone network to output a plurality of groups of initial feature maps at different depths, which are denoted as multi-scale feature layers 、 、 , wherein, 、 、 is the number of channels, 、 、 、 、 、 is the spatial dimension of the feature map, respectively.

[0052] In this step, these feature layers respectively encode multi-granularity and high-correlation information such as global background, personnel posture and equipment structure.

[0053] S203 inputs the multi-scale feature layers into the path enhancement module to enhance the information flow and expression between the multi-scale feature layers , and then inputs the multi-scale feature layers into the multi-scale feature fusion module for bottom-up and top-down information fusion to obtain cross-scale feature representation .

[0054] In this step, as shown in Figure 2 , the embedded image feature semantics and edge information of each layer obtained after multi-scale feature fusion are different, which enhances the expression capability of each layer to different target sizes, structures and edges. The multi-scale feature fusion module can use a pyramid model (FPN) to fuse feature information of different levels, so that fine-grained structures, spatial details and high-level semantics can fully flow and complement in multiple scale feature layers, effectively improving the detection and representation capability of the model to different sizes and complex structure targets, and solving the problem of significant target size change in actual scenes.

[0055] S204 unifies the cross-scale feature representations into a shared feature space to form a total feature expression tensor wherein, is the unified channel dimension, is the spatial dimension at the down-sampling scale.

[0056] In this step, the total feature expression tensor is used to represent the structural and positional attributes of multiple classes of objects present in the image at different scales. The improved YOLO architecture is optimized for unstructured scene features in substation job site images, and has sensitivity to fine-grained information such as small targets (e.g. tools), large targets (e.g. equipment), weak semantic targets (e.g. personnel posture), etc., supporting subsequent cross-modal semantic fusion and target recognition processes.

[0057] S3 decouples the semantic vector of the violation description prompt sentence for text adaptation processing to obtain an enhanced semantic vector;

[0058] S301 performs semantic abstraction extraction on the semantic vector followed by nonlinear projection and compression to obtain semantic features with abstract structure ;

[0059] In this step, a multi-layer perceptron (MLP) or a gate control (such as a gated neural network, GateMechanism) can be used for nonlinear projection and compression of the features.

[0060] S302 performs element multiplication between a task-guided weight vector and the semantic features with abstract structure to reconstruct the semantic features with abstract structure in the semantic dimension, obtaining an enhanced semantic vector , with the calculation formula as follows:

[0061] ,

[0062] wherein, is a feature reconstruction multi-layer perceptron.

[0063] In this step, the prompt sentence vector obtained in step S1 is semantically enhanced, and a decoupled text adaptation module is used to perform task-specific semantic reconstruction of natural language to generate an enhanced semantic vector , which better meets the detection needs of specific violation behaviors and improves the expression ability of fine-grained violation categories.

[0064] ​S4 fuses the multi-scale semantic features and the enhanced semantic vector to obtain a multi-modal feature, and performs open-class violation behavior recognition and positioning through the multi-modal feature;

[0065] S401 fuses the multi-scale semantic features and the enhanced semantic vector through a text-guided gating mechanism and a multi-head cross-attention mechanism to obtain a text-image fusion feature ;

[0066] This step is shown in Figure 3 , and the specific steps are as follows:

[0067] S4011 divides the total feature expression tensor into two parts, which are respectively input into the CBS module to perform basic feature extraction to obtain basic image features as the original image features ;

[0068] In this step, the CBS module is a group of classic convolutional neural network basic units. The CBS in the name is generally an abbreviation of Conv (convolution), BN (Batch Normalization), and S (activation function, such as SiLU, ReLU, or Swish). The input image features or high-dimensional tensors are standardized and enhanced once to lay a foundation for subsequent fusion.

[0069] S4012 sequentially enhances the basic image features through two deformable convolution layers and , and a standard convolution module to obtain first enhanced image features , and the calculation formula is as follows:

[0070] ;

[0071] In this step, deformable convolution can dynamically adjust the convolution receptive field according to the image content, better capture the geometric changes of the target, and improve the perception ability. This layer of operation enhances the flexibility and expressiveness of the original image features.

[0072] S4013 performs an element-wise multiplication operation on the enhanced image features and the basic image features to obtain second enhanced image features , and the calculation formula is as follows:

[0073] ;

[0074] In this step, the semantic sensitive area related to the prompt sentence is emphasized through the gating mechanism, so that the network can automatically adjust, emphasize and weaken the feature channels of different areas, highlight the positions most relevant to the semantic sentence (such as "no safety helmet"), and these positions are amplified by the gating mechanism, and other irrelevant positions are suppressed.

[0075] S4014 calculates the second enhanced image feature with the enhanced semantic vector , and calculates the maximum value of the inner product of the first text vector and each image position . , and generates an attention map through a Sigmoid activation function , and the calculation formula is as follows:

[0076] ,

[0077] ;

[0078] In this step, the Max-Sigmoid interaction mechanism is used to generate the similarity response between the image and the text. Specifically, the maximum value (max) of the score driven by a certain text in the spatial dimension (for example, all image positions) is first taken, and then the Sigmoid activation function is used to normalize the maximum value result to obtain the score, which is then used to generate the attention mask or weighted output. The Sigmoid activation function is one of the most commonly used normalization nonlinear functions in neural networks.

[0079] S4015 forms the final fusion image feature by splicing the second enhanced image feature and the original image feature , and the calculation formula is as follows:

[0080] ;

[0081] In this step, such merging can be compatible with different levels and different granularity feature information, providing diverse representation for subsequent fusion.

[0082] S4016 aligns and fuses the fusion image feature and the text embedding matrix through a multi-head cross attention mechanism to obtain a graph-text fusion feature .

[0083] In this step, the multi-head attention head is used to calculate the cross-modal correlation between the image feature and the text embedding, effectively realizing the precise alignment of the image region and the semantic description, and obtaining the graph-text fusion feature and the aligned features are transmitted to a subsequent target detection module for detection and positioning of the violation behavior.

[0084] S402 performs open-class violation behavior identification and positioning on the image-text fusion features and outputs a detection result.

[0085] S4021 performs regression prediction on the image-text fusion features by a standard convolutional network to output a bounding box position of each candidate region ;

[0086] In this step, the bounding box position of each candidate region describes the spatial position of the target in the image.

[0087] S4022 dynamically generates a category embedding matrix from the image-text fusion features wherein, is the number of candidate categories, is the image-text shared embedding dimension;

[0088] In this step, when performing feature classification identification, the C-class softmax classifier based on fixed convolution kernels in traditional detection methods is abandoned. The traditional detection method can only identify fixed closed-class categories. The weight of each category is dynamically obtained from the category text description (such as “not wearing a safety helmet” and “entering a dangerous area without authorization”) through a text encoder, which can flexibly supplement new categories and does not rely on fixed parameters.

[0089] S4023 flattens the image-text fusion features into a two-dimensional form and then performs a transpose operation, with the calculation formula being as follows:

[0090] ,

[0091] ,

[0092] In this step, in order to efficiently perform matrix calculation, the spatial features are flattened into a two-dimensional form and then transposed, each row of which represents a feature vector of a spatial position in the image, which is used for inner product with the category embedding in the subsequent step.

[0093] S4024 performs matrix inner product calculation on the flattened and transposed image-text fusion features and the category embedding matrix to generate a similarity matrix wherein each element represents the similarity between the i-th category and the j-th category. ​similarity scores between image regions;

[0094] In this step, the essence is the multi-modal fusion of image regions and category semantics, which can find which region in the image is more like the semantics of a certain category.

[0095] S4025 flattens the similarity matrix into a spatial distribution form and performs normalized activation to obtain the semantic response strength of each position to each category;

[0096] In this step, the Sigmoid function can be used for normalized activation, which is suitable for multi-label; the Softmax function can be used for normalized activation, which is suitable for single label. This step is used to identify open category targets.

[0097] S4026 calculates the semantic response strength of each category for the boundary box position of each candidate region in the spatial distribution form and takes the category with the strongest response as the predicted target category and outputs the detection result .

[0098] In this step, represents the target category predicted by the image-text semantic matching mechanism, and the corresponding boundary box position is obtained in step S4021. The image-text fusion feature is input into the decoupled detection head, which includes two sub-branches of position regression and category identification. The position regression branch accurately locates the spatial position of the candidate target boundary box, and the category identification branch fully utilizes the multi-modal image-text fusion feature and the category embedding dynamically generated by the text encoder to realize open category identification for each candidate region. This decoupled detection head design complements the advantages of spatial positioning and semantic classification, especially in the category identification link, which uses spatial distribution form semantic response strength (Label) for regional aggregation decision, achieving accurate identification and positioning of open category violations in substations.

[0099] S5 embeds and caches similar and repeated violation description prompt sentences, and deploys them on edge devices for real-time violation behavior detection.

[0100] S501 receives user input violation description prompt sentences in the inference phase encodes them into semantic vectors and generates a unique key for the semantic vector through hash digest.

[0101] In this step, the inference stage refers to that the model has been trained and enters the actual application and actual operation stage (such as identification / judgment / prediction on a server, front end, edge device, etc.).

[0102] S502 determines whether the current violation description prompt statement exists in the embedding cache pool, if it exists, the semantic vector in the cache is directly reused , and the encoding and semantic enhancement steps are skipped;

[0103] In this step, by directly copying the semantic vector in the cache, repeated calculation can be avoided.

[0104] S503 if the violation description prompt statement is not recorded in the embedding cache pool, natural language encoding and semantic enhancement are performed, and the semantic vector encoded and enhanced is written in the cache pool together with its key for subsequent repeated calls;

[0105] S504 uses the LRU strategy to dynamically manage the memory space, and preferentially retains the recently high-frequency used violation description prompt statement embedding.

[0106] In this step, the embedding cache mechanism uses the least recently used (LRU, Least Recently Used) strategy to dynamically manage the memory space, preferentially retains the recently high-frequency used prompt statement embedding, reduces the risk of cache overflow and invalidation, and ensures that the memory occupation of the edge device is controllable; in combination with the embedding cache mechanism, the vector reuse strategy is maximized in the embedding generation, image-text alignment and detection calculation process, and the response speed is improved. In the inference process, the embedding cache mechanism is introduced, similar or repeated prompt statement embeddings are cached and reused, repeated calculation is reduced, system inference efficiency is improved, and real-time deployment capability of the detection model on the edge computing device is realized.

[0107] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0108] The present application is described in reference to the flowchart and / or block diagram of the method, apparatus (system) and computer program product according to an embodiment of the present application. It is understood that each flow and / or block in the flowchart and / or block diagram, and a combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 The flowchart and / or block diagram block or blocks Figure 1 The flowchart and / or block diagram block or blocks

[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart and / or block diagram block or blocks. Figure 1 The flowchart and / or block diagram block or blocks Figure 1 The flowchart and / or block diagram block or blocks

[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 The flowchart and / or block diagram block or blocks Figure 1 The flowchart and / or block diagram block or blocks

[0111] It should be noted that any references made in the claims to an "apparatus" or "means" should not be construed to cover the corresponding structures only. The phraseology and terminology employed herein are for the purpose of description and should not be regarded as limiting. The use of "including" and "comprising" and variations thereof is meant to encompass the items listed thereafter and equivalents thereof as well as additional items. The terms "connected," "coupled," and "pathway" are used broadly and encompass both direct and indirect connections, couplings and pathways.

[0112] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those skilled in the art once they learn of the basic inventive concepts. Therefore, the appended claims are intended to cover all such variations and modifications as fall within the scope of the application.

[0113] Obviously, various modifications and changes can be made to the present application by those skilled in the art without departing from the spirit and scope of the present application. Accordingly, the present application intends to include all such modifications and changes as fall within the scope of the claims and their equivalents.

[0114] In the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection", "linking", "fixing" and the like should be understood in a broad sense, for example, can be fixed connection, can also be detachable connection, or integral; can be mechanical connection, can also be electrical connection; can be direct connection, can also be indirect connection through intermediate medium, can be internal communication of two elements or interaction relationship of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0115] In the description of the present application, the description of the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in the present application and the features of the different embodiments or examples without contradiction.

Claims

1. A method for detecting violations in substations based on real-time open vocabulary detection, characterized in that, include: Images of substation operations are collected, and violation descriptions are extracted from the operation specifications document. These violation descriptions are then encoded into a warning vector. Specifically: Images of substation operations are collected. Violation descriptions and alerts are extracted from the corresponding work specifications document for each task type and tagged accordingly. ,in, Generate task-related semantic input pairs from a set of violation description prompt statements; For each violation description prompt statement Semantic encoding is performed to obtain corresponding semantic vectors, and the set of violation description prompts is then processed using a pre-trained multimodal language model. Encode to generate semantic vectors ,in, Representing a d-dimensional real vector space, mapping and encoding the set of object labels into a label embedding matrix, and combining semantic vectors and label embedding matrices to form a task-related subset of labels; An improved YOLO architecture was constructed by using a backbone extraction network, a path enhancement module, and a multi-scale feature fusion module. The on-site operation images were input into the feature extraction network of the improved YOLO architecture to obtain multi-scale image semantic features. For the detection task, the semantic vector of the violation description and prompt statement is subjected to decoupled text adaptation processing to obtain an enhanced semantic vector, specifically: semantic vectors After semantic abstraction and extraction, nonlinear projection and compression are performed to obtain semantic features with abstract structures. ; Guide the weight vector through a task semantic features with abstract structure Element-wise multiplication, combined with prior hints from the violation detection scenario, is applied to semantic features with abstract structures. Reconstruction along the semantic dimension yields an enhanced semantic vector. The calculation formula is as follows: , in, Reconstructing a multilayer perceptron for features; Multi-scale semantic features and enhanced semantic vectors are fused and aligned to obtain multi-modal features. These multi-modal features are then used to identify and locate open-category violations. Specifically: Multi-scale semantic features and enhanced semantic vectors are fused using a text-guided gating mechanism and a multi-head cross-attention mechanism to obtain image-text fusion features. ; Image-text fusion features Identify and locate open-category violations, and output the detection results; Similar and repetitive violation descriptions are embedded and cached, and then deployed on edge devices for real-time violation detection.

2. The substation violation detection method based on real-time open vocabulary detection according to claim 1, characterized in that, The process involves acquiring images of substation operations, extracting violation descriptions from the corresponding work specifications document for each task type, tagging the corresponding objects, and generating task-related semantic input pairs, including: Video capture of substation operation scenes ,in, and These represent the height and width of the image, respectively, and 3 indicates a three-channel image; The violation description prompts are processed into verb-object structures, and their semantics are aligned with the violation behavior dictionary. Define the object label set according to the actual task scenario. Each of the tags Indicates category terms and / or specific objects related to the scene; Description and prompts for traffic violations With object label collection Perform consistency verification, remove semantic duplicates and / or category redundancies, and generate task-related semantic input pairs. .

3. The substation violation detection method based on real-time open vocabulary detection according to claim 1, characterized in that, The description and prompt statements for each violation Semantic encoding is performed to obtain the corresponding semantic vector. The object label set is mapped and encoded into a label embedding matrix. The semantic vector and the label embedding matrix are jointly used to form a task-related label subset, including: If the set of violation description prompt statements If it is a compound statement, it will be semantically segmented into atomic prompts. Then the atomic hints Encoded as And they are fused into a unified representation through an attention mechanism; Collection of object labels Each tag item in Mapped to the corresponding label vector A tag embedding matrix is ​​formed through the encoder. ,in, Represents the set of real matrices with n rows and d columns; semantic vectors and tag embedding matrix Joint modeling is performed using semantic similarity calculation. The tag set is reordered or filtered to form a subset of task-related tags. ; If the set of violation description prompt statements If it is an empty set, then the set of categories constructed during the training phase will be used. As a set of target categories.

4. The substation violation detection method based on real-time open vocabulary detection according to claim 1, characterized in that, The process of inputting on-site operation images into a feature extraction network with an improved YOLO architecture to obtain multi-scale image semantic features includes: The backbone network performs multi-level convolution and spatial downsampling on the field operation images, outputting multiple sets of initial feature maps at different depths, which are denoted as multi-scale feature layers. , , ,in, , , For the number of channels, , , , , , These represent the spatial dimensions of the feature map; Multi-scale feature layers Input path enhancement module enhances multi-scale feature layers. The information flow and expression between them are then input into the multi-scale feature fusion module for the multi-scale feature layer. By fusing bottom-up and top-down information, cross-scale feature representations are obtained. ; Representing cross-scale features A unified mapping is performed to a shared feature space, forming a total feature representation tensor. ,in, To unify channel dimensions, This represents the spatial dimension at the downsampling scale.

5. The substation violation detection method based on real-time open vocabulary detection according to claim 1, characterized in that, The process of fusing multi-scale semantic features and enhanced semantic vectors is achieved through a text-guided gating mechanism and a multi-head cross-attention mechanism to obtain image-text fusion features. include: The total feature expression tensor Divided into two parts, of which, To unify channel dimensions, The spatial dimensions at the downsampling scale are used as inputs to the CBS module for basic feature extraction, yielding the basic image features. As original image features ; Basic image features Passing through two deformable convolutional layers in sequence and A standard convolutional module is used for enhancement processing to obtain the first enhanced image features. The calculation formula is as follows: ; Enhanced image features With basic image features Perform element-wise multiplication to obtain the second enhanced image features. The calculation formula is as follows: ; The second enhanced image feature With enhanced semantic vectors Interact and calculate the first Each text vector and each image location The maximum value of the inner product Attention maps are generated using the Sigmoid activation function. The calculation formula is as follows: , ; By enhancing the second image features Features of the original image The images are then stitched together to form the final fused image features. The calculation formula is as follows: ; Fusing image features and text embedding matrix Image-text fusion features are obtained through alignment and fusion using a multi-head cross-attention mechanism. .

6. The substation violation detection method based on real-time open vocabulary detection according to claim 1, characterized in that, The text-image fusion feature The system identifies and locates open-category violations, and outputs detection results including: Image-text fusion features are applied using a standard convolutional network. Perform regression prediction and output the bounding box location of each candidate region. ; Image and text fusion features Dynamically generate category embedding matrix using text encoder ,in, The number of candidate categories, Embedded dimensions for image and text sharing; Image and text fusion features Flattened into a two-dimensional form Then perform the transpose operation, and the calculation formula is as follows: , , The flattened and transposed image-text fusion feature With category embedding matrix Perform matrix inner product calculation to generate a similarity matrix. , where each element Indicates the first Class and the Similarity scores between image regions; Similarity matrix Flattened into a spatial distribution shape Then, normalized activation is performed to obtain the semantic response strength of each position for each category; For the bounding box location of each candidate region In spatial distribution Statistically analyze the semantic response intensity of each category of the bounding box, and select the category with the strongest response as the predicted target category. Output detection results .

7. The substation violation detection method based on real-time open vocabulary detection according to claim 1, characterized in that, The step of embedding and caching similar and repetitive violation description prompts and deploying them on edge devices for real-time violation detection includes: During the reasoning phase, a set of violation description prompts input by the user is received. Encode it into a semantic vector And generate a unique key by hashing the semantic vector. ; Determine if the current violation description message already has a corresponding embedded record in the embedding cache pool. If it already exists, the enhanced semantic vector in the cache will be reused directly. Skip the encoding and semantic enhancement steps; If the set of violation description prompt statements If not recorded in the embedded cache pool, natural language encoding and semantic enhancement are performed, and the encoded and enhanced semantic vectors are converted into semantic vectors. with its key They are written together into the cache pool for subsequent repeated calls; The LRU strategy is used to dynamically manage memory space, prioritizing the retention of frequently used violation description and prompt statements.

Citation Information

Patent Citations

  • Market illegal behavior detection method based on multi-modal image-text fusion

    CN120145098A