Electric energy metering equipment wiring standard detection method and device, computer equipment and storage medium
By using the combination of local window attention module, global interaction module and sparse attention mechanism in the grid wiring specification detection, the problems of inefficient and insufficient accuracy of traditional detection methods are solved, and more efficient and accurate wiring specification detection is achieved.
Patent Information
- Application Number
- CN202510256264.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-24
AI Technical Summary
The traditional grid wiring specification detection method is inefficient and easily affected by manual subjective judgment, resulting in insufficient detection accuracy, especially in complex scenarios, it is difficult to take into account the relationship between the local details of the image and the global topological structure, resulting in missed inspection of subtle errors.
The detection method based on the local window attention module and the global interaction module is adopted, combined with the sparse attention mechanism, the wiring image and wiring specification text characteristics of the electrical energy metering equipment are obtained, the characteristics are fused through the gate weighting mechanism, and the wiring detection results are determined using the pre-constructed classification head.
The accuracy of wiring specification detection is improved, and the local details and global information of the image can be taken into account. The detection dimension is increased through text features, which significantly improves the detection accuracy of wiring specification detection of power metering equipment.
Smart Images

Figure CN120198376A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent operation and maintenance of power systems, and in particular, to a method, device, computer device, storage medium, and computer program product for detecting the wiring specification of electric energy metering equipment. Background Art
[0002] Whether the wiring of electric energy metering equipment (such as electric energy meters, junction boxes, and terminals, etc.) is standardized can directly relate to the safety and stability of the power grid. With the continuous expansion and complexity of the scale of the power system, traditional wiring specification detection methods have been difficult to meet the requirements of modern power grids for efficient and accurate operation and maintenance. Traditional detection methods mainly rely on manual inspections. This method is not only inefficient but also easily affected by the experience and subjective judgment of inspection personnel, resulting in errors in the detection results.
[0003] In related technologies, a detection method based on image recognition can be used for wiring specification detection and introduced into the power system. However, the detection method based on image recognition often performs poorly when dealing with complex scenarios. For example, in the face of actual application scenarios such as light changes and equipment occlusion, it is difficult to balance the relationship between the local details and the global topological structure of the image, resulting in missed detection of subtle errors and insufficient detection accuracy for the wiring specification detection of electric energy metering equipment. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for detecting the wiring specification of electric energy metering equipment that can improve the detection accuracy of wiring specifications for the above technical problems.
[0005] In a first aspect, the present application provides a method for detecting the wiring specification of electric energy metering equipment. The method includes:
[0006] Obtain a wiring image to be detected corresponding to the electric energy metering equipment, and a wiring specification text corresponding to the wiring image to be detected;
[0007] Based on a local window attention module and a global interaction module, determine the image features corresponding to the wiring image to be detected; based on a sparse attention mechanism, determine the text features corresponding to the wiring specification text; the local window attention module is used to divide the wiring image to be detected into multiple windows and determine the attention weights of each window; the global interaction module is used to fuse the image information within the global range through the attention weights of each window to obtain image features;
[0008] Map the image features and the text features to a unified dimensional space to obtain a first mapping value corresponding to the image features and a second mapping value corresponding to the text features;
[0009] Fuse the first mapping value and the second mapping value according to the gating weighting mechanism to obtain a fused feature, and determine the wiring detection result corresponding to the fused feature based on a pre-constructed classification head.
[0010] In one embodiment, the image feature is a multi-scale image feature; the determining the image feature corresponding to the wiring image to be detected based on the local window attention module and the global interaction module includes:
[0011] Based on the local window attention module and the multi-scale configuration information, divide the wiring image to be detected into a plurality of non-overlapping windows, and determine the local attention features corresponding to each window; the local attention feature is determined by weighted summation of the attention weights of the window and the window features.
[0012] Determine the local attention features corresponding to the plurality of non-overlapping windows as single-scale local features, and obtain a plurality of single-scale local features with different scales; the multi-scale configuration information includes the window sizes corresponding to each single-scale local feature.
[0013] Based on the global interaction module, determine the global feature of the wiring image to be detected.
[0014] Fuse each single-scale local feature with the global feature respectively to obtain a multi-scale image feature.
[0015] In one embodiment, the determining the text feature corresponding to the wiring specification text based on the sparse attention mechanism includes:
[0016] Determine the fixed window corresponding to each token in the wiring specification text, and determine the window text feature determined by a plurality of tokens included in the fixed window.
[0017] Obtain the target tokens with global attention, and determine the global text features determined by each target token.
[0018] Fuse each window text feature with the global text feature respectively to obtain the text feature corresponding to the wiring specification text.
[0019] In one embodiment, the mapping the image feature and the text feature to a unified dimensional space to obtain the first mapping value corresponding to the image feature and the second mapping value corresponding to the text feature includes:
[0020] Determine the first linear projection layer corresponding to the image feature and the second linear projection layer corresponding to the text feature; the first linear projection layer is used to map the input dimension corresponding to the image feature to a preset dimension; the second linear projection layer is used to map the input dimension corresponding to the text feature to the preset dimension;
[0021] Based on the first linear projection layer, map the image feature to the dimension space corresponding to the preset dimension to obtain the first mapping value corresponding to the image feature;
[0022] Based on the second linear projection layer, map the text feature to the dimension space corresponding to the preset dimension to obtain the second mapping value corresponding to the text feature.
[0023] In one embodiment, the step of fusing the first mapping value and the second mapping value according to the gating weighting mechanism to obtain a fused feature, and determining the wiring detection result corresponding to the fused feature based on a pre-constructed classification head includes:
[0024] Based on the gating weight matrix in the gating weighting mechanism, perform weighted summation on the first mapping value and the second mapping value to obtain a fused feature;
[0025] Input the fused feature into a classification head constructed based on Transformer to obtain the probability values of each wiring detection category output by the classification head;
[0026] Based on the probability values of each wiring detection category, determine the wiring detection result corresponding to the fused feature.
[0027] In one embodiment, the method further includes:
[0028] Determine the contrast loss function corresponding to the image feature and the text feature, and the formula of the contrast loss function is as follows:
[0029]
[0030]
[0031] Wherein, is the similarity score, is the transpose of the first mapping value corresponding to the image feature, is the second mapping value corresponding to the text feature, is the temperature coefficient, N is the number of negative samples within a batch, is the k-th text negative sample that does not match the first mapping value, is the k-th image negative sample that does not match the second mapping value;
[0032] Determine the cross-entropy loss as the classification loss function;
[0033] Weighted sum the contrast loss function and the classification loss function to obtain a combined loss function; the combined loss function is used to train the model corresponding to the classification head.
[0034] In one embodiment, the method further includes:
[0035] For the contrast loss function, the formula for adjusting the temperature coefficient is as follows:
[0036]
[0037] where is the initial temperature value, α is the adjustment coefficient, is the similarity score.
[0038] In a second aspect, the present application also provides a device for detecting the wiring specification of an electric energy metering device. The device includes:
[0039] A data acquisition module, configured to acquire a wiring image to be detected corresponding to an electric energy metering device, and a wiring specification text corresponding to the wiring image to be detected;
[0040] A feature determination module, configured to determine an image feature corresponding to the wiring image to be detected based on a local window attention module and a global interaction module; determine a text feature corresponding to the wiring specification text based on a sparse attention mechanism; the local window attention module is configured to divide the wiring image to be detected into multiple windows and determine the attention weights of each window; the global interaction module is configured to fuse image information within a global range through the attention weights of each window to obtain an image feature;
[0041] A feature mapping module, configured to map the image feature and the text feature to a unified dimensional space to obtain a first mapping value corresponding to the image feature and a second mapping value corresponding to the text feature;
[0042] A feature detection module, configured to fuse the first mapping value and the second mapping value according to a gated weighting mechanism to obtain a fused feature, and determine a wiring detection result corresponding to the fused feature based on a pre-constructed classification head.
[0043] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect are implemented.
[0044] Fourthly, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0045] Fifthly, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0046] For the above power metering device wiring specification detection method, device, computer device, storage medium and computer program product, by obtaining the wiring image to be detected corresponding to the power metering device and the wiring specification text corresponding to the wiring image to be detected, and based on the local window attention module and the global interaction module, the image features corresponding to the wiring image to be detected are determined, and based on the sparse attention mechanism, the text features corresponding to the wiring specification text are determined. The local window attention module can divide the wiring image to be detected into multiple windows and determine the attention weights of each window, while the global interaction module can fuse the image information within the global range through the attention weights of each window to obtain image features. Then, the image features and text features are mapped to a unified dimensional space to obtain the first mapping value corresponding to the image features and the second mapping value corresponding to the text features. Finally, the first mapping value and the second mapping value are fused according to the gated weighting mechanism to obtain the fused features, and based on the pre-constructed classification head, the wiring detection result corresponding to the fused features is determined and used as the wiring detection result of the power metering device to determine whether the power metering device is wired incorrectly. The embodiments of the present application can respectively extract the local features and global features of the wiring image to be detected and the wiring specification text to obtain image features and text features, and fully fuse the image features and text features as multi-modal features for classification, which can take into account both the local details and global information of the image, and can further increase the feature dimension for wiring specification detection through text features, thereby improving the detection accuracy of the wiring specification detection of the metering device. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 It is an application environment diagram of the power metering device wiring specification detection method in an embodiment;
[0049] Figure 2Schematic flowchart of a method for detecting wiring specifications of an electric energy metering device in an embodiment;
[0050] Figure 3 Schematic flowchart of a method for determining image features of a wiring image to be detected in an embodiment;
[0051] Figure 4 Schematic flowchart of a method for determining text features of a wiring image to be detected in an embodiment;
[0052] Figure 5 Schematic flowchart of a method for detecting wiring specifications of an electric energy metering device in another embodiment;
[0053] Figure 6 Block diagram of the structure of a device for detecting wiring specifications of an electric energy metering device in an embodiment;
[0054] Figure 7 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0055] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0056] The method for detecting wiring specifications of an electric energy metering device provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 can obtain the to-be-detected wiring images of each to-be-detected power metering device, as well as the wiring specification texts corresponding to each to-be-detected power metering device. For the same type of power metering device, the server 104 can receive the to-be-detected wiring image and the wiring specification text from the terminal 102. The server 104 can determine the image features corresponding to the to-be-detected wiring image through a trained model including a local window attention module and a global interaction module. The server 104 can determine the text features corresponding to the wiring specification text through a trained model based on a sparse attention mechanism. The server 104 can map the image features and text features of the corresponding power metering device to a unified dimensional space, obtain the first mapping value of the image features in this dimensional space, and the second mapping value of the text features in this dimensional space. The server 104 can fuse the first mapping value and the second mapping value through a pre-configured gating weighting mechanism to obtain a fused feature. Finally, the server 104 calls a classification head to determine the classification result corresponding to the fused feature and use it as the wiring detection result. For different types of power metering devices, corresponding models can be pre-configured or trained in advance, such as a model including a local window attention module and a global interaction module, a model including a sparse attention mechanism, etc.
[0057] The data storage system can store the data that the server 104 needs to process, such as the to-be-detected wiring images corresponding to each to-be-detected power metering device, and the wiring specification texts corresponding to each power metering device. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers.
[0058] Among them, the terminal 102 can be but is not limited to various terminals such as personal computers, laptop computers, smart phones, and tablet computers. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0059] In an exemplary embodiment, as Figure 2 shown, a method for detecting the wiring specification of a power metering device is provided. Taking the method applied to the Figure 1 server 104 in it as an example for description, it includes the following steps S202 to step S208. Among them:
[0060] Step S202, obtain the to-be-detected wiring image corresponding to the power metering device, and the wiring specification text corresponding to the to-be-detected wiring image.
[0061] Among them, the electric energy metering device is used to measure the quantity of electric energy. The electric energy metering device is an instrument for measuring and calculating the electricity quantity, and can be used to measure the power generation, plant electricity consumption, power supply, electricity sales, etc. Common electric energy metering devices include electric energy meters, instrument transformers, electric energy metering cabinets, etc. The to-be-detected wiring image refers to the image obtained by photographing the electric energy metering device under different lights and different photographing angles through a photographing device. The wiring specification text is a document that clearly stipulates and requires various aspects such as wiring operations, processes, materials, etc., including operation manuals, equipment parameter tables, etc.
[0062] Specifically, the server can obtain the to-be-detected wiring images corresponding to each to-be-detected electric energy metering device from each terminal, and at the same time obtain the wiring specification text corresponding to the to-be-detected wiring image. In one example, the server can determine the device model or device type of the to-be-detected electric energy metering device, and determine the wiring specification text corresponding to the electric energy metering device according to the device model or device type.
[0063] Step S204, based on the local window attention module and the global interaction module, determine the image features corresponding to the to-be-detected wiring image; based on the sparse attention mechanism, determine the text features corresponding to the wiring specification text.
[0064] Among them, the local window attention module is used to divide the to-be-detected wiring image into multiple windows and determine the attention weights of each window; the global interaction module is used to fuse the image information within the global range through the attention weights of each window to obtain image features. The sparse attention mechanism is an efficient attention mechanism used to solve the problem of excessive computational complexity of the traditional attention mechanism when processing long sequence data. The sparse attention mechanism adopts a dynamic hierarchical sparse strategy, combines coarse-grained Token compression and fine-grained Token selection, enables the model to select key information, only performs attention calculations on some important positions, reduces unnecessary calculations, and thus significantly improves the computational efficiency while maintaining the accuracy of information processing.
[0065] Specifically, the server can pre-train an image feature extraction model including a local window attention module and a global interaction module, and input the image features corresponding to the to-be-detected wiring image into the image feature extraction model. The image feature extraction model can divide the to-be-detected wiring image through the local window attention module to obtain multiple windows, determine the attention weights of each window, and determine the features corresponding to each window based on the attention weights and the feature information included in the windows. In addition, the server can fuse the features corresponding to each window through the global interaction module and the attention weights of each window to obtain the image features corresponding to the to-be-detected wiring image.
[0066] In addition, the server can calculate the attention for the text at some important positions in the wiring specification text based on the sparse attention mechanism, obtain the features corresponding to each text, and determine the text feature corresponding to the wiring specification text through the attention weights corresponding to each text and the features corresponding to each text.
[0067] Step S206: Map the image feature and the text feature to a unified dimensional space to obtain a first mapped value corresponding to the image feature and a second mapped value corresponding to the text feature.
[0068] Among them, the dimensional space of the image feature itself is different from that of the text feature. The server can map the image feature and the text feature to the same dimensional space through mapping.
[0069] Specifically, the server can map the image feature to the unified dimensional space through the first mapping strategy to obtain the first mapped value of the image feature in this dimensional space. At the same time, the server can map the image feature to the unified dimensional space through the second mapping strategy to obtain the second mapped value of the text feature in this dimensional space.
[0070] Step S208: Fuse the first mapped value and the second mapped value according to the gating weighting mechanism to obtain a fused feature, and determine the wiring detection result corresponding to the fused feature based on the pre-constructed classification head.
[0071] Among them, the gating weighting mechanism can assign dynamic weights to different modality data to control the contribution degree of each modality data in the fusion process. In the embodiments of the present application, the modality data includes image data and text data, that is, the wiring image to be detected and the wiring specification text. The gating weighting mechanism can include a pre-configured gating weight matrix. The classification head is an important part of the model architecture, usually located at the end, and is mainly used to implement the classification task of the input data.
[0072] Specifically, the server can determine the weight values corresponding to the first mapped value and the second mapped value through the gating weight matrix in the gating weighting mechanism, and perform weighted summation on the first mapped value and the second mapped value through the weight values to obtain the fused feature. The server can input the fused feature into the pre-constructed classification head. After the classification head processes the fused feature, the classification result is obtained. For example, the classification head can output the probabilities of each classification result, and select the classification result with the largest probability value and the probability value greater than the preset value as the wiring detection result. The pre-constructed classification head can be obtained by training based on the wiring images and wiring specification texts corresponding to the same type of power metering equipment.
[0073] In the above method for detecting the wiring specification of the electric energy metering device, by obtaining the wiring image to be detected corresponding to the electric energy metering device and the wiring specification text corresponding to the wiring image to be detected, and based on the local window attention module and the global interaction module, the image features corresponding to the wiring image to be detected are determined, and based on the sparse attention mechanism, the text features corresponding to the wiring specification text are determined. The local window attention module can divide the wiring image to be detected into multiple windows and determine the attention weights of each window, while the global interaction module can fuse the image information within the global range through the attention weights of each window to obtain image features. Then, the image features and the text features are mapped to a unified dimensional space to obtain the first mapping value corresponding to the image features and the second mapping value corresponding to the text features. Finally, the first mapping value and the second mapping value are fused according to the gated weighting mechanism to obtain the fused features, and based on the pre-constructed classification head, the wiring detection result corresponding to the fused features is determined and used as the wiring detection result of the electric energy metering device to determine whether the wiring of the electric energy metering device is incorrect. The embodiments of the present application can extract the local features and global features of the wiring image to be detected and the wiring specification text respectively to obtain image features and text features, and fully fuse the image features and the text features as multi-modal features for classification, which can take into account the local details and global information of the image, and can further increase the feature dimension for wiring specification detection through the text features, thereby improving the detection accuracy of the wiring specification detection of the metering device.
[0074] In an exemplary embodiment, the image features are multi-scale image features. As Figure 3 shown, the specific implementation process of the step "Based on the local window attention module and the global interaction module, determine the image features corresponding to the wiring image to be detected" includes steps S302 to S308. Among them:
[0075] Step S302, based on the local window attention module and the multi-scale configuration information, divide the wiring image to be detected into multiple non-overlapping windows and determine the local attention features corresponding to each window.
[0076] Among them, the local attention features are determined by weighted summation of the attention weights of the windows and the window features. The multi-scale configuration information includes the configuration information for dividing the wiring image to be detected into different scales, and the configuration information may include the window size of each scale. The window features are the feature information included in the windows.
[0077] Specifically, the server can input the wiring image to be detected into the local window attention module, and perform multiple rounds of processing on the wiring image to be detected according to the multi-scale configuration information. In each round of processing, the wiring image to be detected can be divided into multiple non-overlapping windows corresponding to different scales. For example, in the image, the image can be segmented into small regions of a fixed size, and each small region is a window. For each window, the server can perform self-attention calculation inside each window. The server regards the features within the window as a sequence, calculates the dependence relationship between each position within the window through the standard self-attention mechanism, and captures the local feature information within the window. The specific process includes: mapping the features within the window into query (Q), key (K), and value (V) vectors, calculating the similarity score between Q and K, normalizing the similarity score to obtain the attention weight, and then multiplying the attention weight by V and performing weighted summation to obtain the local attention feature within the window.
[0078] Step S304: Determine the local attention features corresponding to multiple non-overlapping windows as single-scale local features, and obtain multiple single-scale local features of different scales.
[0079] Among them, the multi-scale configuration information includes the window sizes corresponding to each single-scale local feature.
[0080] Specifically, the server can merge the local attention features corresponding to multiple non-overlapping windows to obtain single-scale local features. Through the multi-scale configuration information, the server can process the wiring image to be detected through the local window attention module to obtain multiple single-scale local features of different scales. For the multi-scale configuration information, when generating single-scale local features in the early stage, the window may be small and the image blocks are divided finely, mainly capturing detailed information. Subsequently, as the stacking progresses, the window gradually becomes larger and the image block division becomes coarser, capturing more global semantic information.
[0081] Step S306: Based on the global interaction module, determine the global feature of the wiring image to be detected.
[0082] Specifically, the server can divide the wiring image to be detected into a fixed number of preset image blocks through the global interaction module. It should be understood that each image block is not of a fixed size, but is divided within the entire image range. The global interaction module performs self-attention at the corresponding positions of each image block to obtain global context information. Then, the server calculates the attention weights between the corresponding positions of different image blocks through the global interaction module, and fuses the information within the global range to obtain the global feature of the wiring image to be detected.
[0083] Step S308: Fuse each single-scale local feature with the global feature respectively to obtain multi-scale image features.
[0084] Specifically, the server can fuse multiple single-scale local features extracted by the local window attention module and the global features extracted by the global interaction module. In one example, generally, methods such as residual connection are used to add the local features and the global features or perform other forms of combination to obtain multi-scale image features, which can make the multi-scale image features contain rich local details and fuse global context information.
[0085] In this embodiment, multiple single-scale local features are determined through the local window attention module and multi-scale configuration information, and global features are determined through the global interaction module. Then, the multiple single-scale local features and the global features are fused to obtain multi-scale image features, which can improve the richness and detail level of the multi-scale image features.
[0086] In an exemplary embodiment, as Figure 4 shown, the specific implementation process of the step "determine the text features corresponding to the wiring specification text based on the sparse attention mechanism" includes steps S402 to S406. Among them:
[0087] Step S402, determine the fixed window corresponding to each token in the wiring specification text, and determine the window text features determined by multiple tokens included in the fixed window.
[0088] Specifically, the server performs a tokenization operation on the input wiring specification document, splits the text into multiple tokens, and generates corresponding token embedding vectors for each token. At the same time, positional embeddings are added to the token embedding vectors. The positional embeddings are used to represent the position information of the tokens in the text sequence, enabling the model to perceive the order relationship between words.
[0089] Based on this, the server can define a fixed-size fixed window for each token. The token only calculates the attention scores with other tokens within the window, and determines the window text features corresponding to the token within the window based on the attention scores. For example, the server can perform a weighted sum of the attention scores and the text features corresponding to other tokens within the window to obtain the window text features.
[0090] In one example, for a sequence of length n, each token may only interact with the w tokens before and after it (w is the window size). This method enables each token to capture the local context information around it and reduces the computational complexity.
[0091] Optionally, the token embedding vectors can be a semantic keyword library constructed through wiring specification text descriptions such as operation manuals and device parameter tables. The semantic keyword library can contain initial representations generated through pre-trained token embeddings, that is, the initial representations can be the token embedding vectors of each token.
[0092] Step S404: Obtain the target tokens with global attention and determine the global text features determined by each target token.
[0093] Specifically, the server can configure the target tokens with global attention, interact the target tokens with all other tokens except the target tokens, and obtain the global context information. Then, the server determines the global text features of each target token in the wiring specification text based on the attention scores corresponding to each token in the wiring specification text and the global attention of the target tokens. For example, the global text features corresponding to the target tokens are obtained by weighted summation of the attention scores corresponding to each token and the global attention of the target tokens.
[0094] Step S406: Fuse each window text feature with the global text feature to obtain the text feature corresponding to the wiring specification text.
[0095] Specifically, for the window text feature corresponding to each fixed window, the server can fuse the window text feature with the global text feature to obtain the text feature of the token corresponding to the window. After traversing all the fixed windows, the text features corresponding to each token in the wiring specification text can be obtained.
[0096] In this embodiment, fixed windows are created for each token in the wiring specification text through the sparse attention mechanism, and the window text features corresponding to the fixed windows are determined. The global text features corresponding to the target tokens are obtained through the target tokens with global attention. By fusing each window text feature with the global text feature, the text features corresponding to each token in the wiring specification text can be obtained, which can improve the richness and detail of the text features.
[0097] In an exemplary embodiment, the specific implementation process of the step "map the image feature and the text feature to a unified dimensional space to obtain the first mapping value corresponding to the image feature and the second mapping value corresponding to the text feature" includes:
[0098] Determine the first linear projection layer corresponding to the image feature and the second linear projection layer corresponding to the text feature; based on the first linear projection layer, map the image feature to the dimensional space corresponding to the preset dimension to obtain the first mapping value corresponding to the image feature; based on the second linear projection layer, map the text feature to the dimensional space corresponding to the preset dimension to obtain the second mapping value corresponding to the text feature.
[0099] Among them, the first linear projection layer is used to map the input dimension corresponding to the image feature to the preset dimension; the second linear projection layer is used to map the input dimension corresponding to the text feature to the preset dimension.
[0100] Specifically, before starting the mapping, the server can obtain the initial dimensions of the image features and the text features. For example, assume that the input dimension of the image features is D image , and the input dimension of the text features is D text , D image and D text are not the same. The server can define a first linear projection layer, and the input dimension of the first linear projection layer can be D image , and the output dimension is a preset dimension D target . The server can define a second linear projection layer, and the input dimension of the second linear projection layer can be D text , and the output dimension is a preset dimension D target . The server can input the image features into the first linear projection layer to obtain the first mapping value of the image features in the dimension space of the preset dimension. The server can input the text features into the second linear projection layer to obtain the second mapping value of the text features in the dimension space of the preset dimension.
[0101] In this embodiment, through two linear projection layers, the image features and text features of different dimensions can be projected into the same dimension space to obtain the first mapping value and the second mapping value, which can improve the accuracy of determining the first mapping value and the second mapping value.
[0102] In an exemplary embodiment, the specific implementation process of the step "fuse the first mapping value and the second mapping value according to the gating weighting mechanism to obtain a fused feature, and determine the wiring detection result corresponding to the fused feature based on the pre-constructed classification head" includes:
[0103] Based on the gating weight matrix in the gating weighting mechanism, perform weighted summation on the first mapping value and the second mapping value to obtain a fused feature; input the fused feature into the classification head constructed based on Transformer to obtain the probability values of each wiring detection category output by the classification head; based on the probability values of each wiring detection category, determine the wiring detection result corresponding to the fused feature.
[0104] Among them, Transformer adopts an encoder-decoder architecture, which is a sequence-to-sequence model structure. The encoder is responsible for encoding the input sequence into a fixed-length vector representation, and the decoder generates the next output based on the output of the encoder and the previously generated output sequence.
[0105] Specifically, the server can determine the weight values corresponding to the first mapping value and the second mapping value through the gating weight matrix in the gating weighting mechanism, and perform weighted summation on the first mapping value and the second mapping value through the weight values to obtain a fused feature.
[0106] In an example, the formula for determining the fused feature is as follows:
[0107]
[0108] Among them, σ is the Sigmoid function, and ⊙ represents element-wise multiplication. is the gating weight matrix. is the transpose of the first mapping value corresponding to the image feature. is the second mapping value corresponding to the text feature. is the bias.
[0109] Then, the server can input the fused feature into a classification head constructed based on Transformer. The classification head can output the probabilities of each category for the fused feature, and the server can obtain the probability values of each wiring detection category. For example, the formula for the wiring detection category probability output by the Transformer-based classification head is as follows:
[0110]
[0111] Among them, C is the number of categories. is the fused feature. is the bias.
[0112] Finally, the server can determine the wiring detection result corresponding to the fused feature based on the probability values of each wiring detection category. In one example, determine the probability value of the wiring detection category with the largest probability value among the probability values of different wiring detection categories, and determine whether this probability value is greater than a preset value. If this probability value is greater than the preset value, then determine that the wiring detection result is the result corresponding to this category. For example, the output wiring detection categories include correct wiring, wrong wiring, and loose connection, and the probability of correct wiring is 0.7, the probability of wrong wiring is 0.2, the probability of loose connection is 0.1, and the preset value is 0.6, then determine that the wiring detection result is correct wiring.
[0113] In this embodiment, by weighted summing the first mapping value and the second mapping value through the gating weight matrix to obtain the fused feature, and determining the probability value corresponding to the fused feature through the classification head, and determining the wiring detection result according to the probability values of each wiring detection category, the accuracy of wiring specification detection can be improved.
[0114] In an exemplary embodiment, the method further includes:
[0115] Determine the contrast loss function corresponding to the image feature and the text feature. The formula for the contrast loss function is as follows:
[0116]
[0117]
[0118] Among them, is the similarity score. is the transpose of the first mapping value corresponding to the image feature, is the second mapping value corresponding to the text feature, is the temperature coefficient, and N is the number of negative samples within a batch, is the k-th text negative sample that does not match the first mapping value, is the k-th image negative sample that does not match the second mapping value.
[0119] Specifically, the contrast loss function is a commonly used loss function in machine learning, which is used to measure the similarity or difference between two samples. For example, it measures the difference between image features and text features. The contrast loss function can be used to train a model to make the corresponding text features and image features closer.
[0120] After that, the server can determine the cross-entropy loss as the classification loss function; weight and sum the contrast loss function and the classification loss function to obtain the joint loss function; the joint loss function is used to train the model corresponding to the classification head.
[0121] Specifically, the cross-entropy loss can be:
[0122]
[0123] where C is the number of categories, and y c is an indicator variable. If the true category of the sample is the c-th category, then y c = 1, and for other categories, y c is 0. p c is the probability that the model predicts the sample belongs to the -th category.
[0124] Weight and sum the contrast loss function and the classification loss function to obtain the formula for the joint loss function as follows:
[0125]
[0126] In this embodiment, by weighting and summing the contrast loss function and the classification loss function to obtain the joint loss function, it can make the corresponding multiple models in this application embodiment be trained based on the joint loss function to obtain a trained model, improving the prediction accuracy of the classification head.
[0127] In an exemplary embodiment, the method further includes:
[0128] For the contrast loss function, the formula for adjusting the temperature coefficient is as follows:
[0129]
[0130] where, is the initial temperature value, which is a preset fixed value and serves as the starting benchmark for dynamic adjustment. α is the adjustment coefficient, which is used to control the influence degree of sample difficulty on the adjustment of the temperature coefficient. is the similarity score, and N is the number of samples within a batch. represents the final dynamically obtained temperature coefficient, which will change dynamically according to the sample situation within a batch. represents the feature information of the i-th sample within a batch, such as image features. represents the target information of the i-th sample within a batch, such as text features.
[0131] such as Figure 5 As shown below, a specific embodiment is combined to describe in detail the specific execution process of the above-mentioned wiring specification detection of the electric energy metering device, including the following steps:
[0132] Step S1: Multimodal data construction and preprocessing.
[0133] 1. Image data acquisition and enhancement:
[0134] Obtain the wiring image dataset of the electric energy metering device (electric energy meter, wiring box, terminal), covering different lighting conditions, angles, and occlusion scenarios. Perform size normalization (adjust to ), Gaussian denoising (kernel size ), and data enhancement (random rotation , brightness adjustment ) on the original image.
[0135] 2. Text data annotation and cleaning:
[0136] Collect the wiring specification text descriptions corresponding to the images (such as operation manuals, device parameter tables), perform word segmentation and remove stop words, and construct a semantic keyword library. The text is vectorized using pre-trained word embeddings (such as Word2Vec) to generate initial representations , where = 300 is the embedding dimension.
[0137] Step S2: Multimodal feature extraction.
[0138] 1. Image feature extraction (MaxViT model):
[0139] The improved Vision Transformer (MaxViT) extracts multi-scale features by alternately stacking local window attention and global interaction blocks.
[0140] Among them, local window attention: The image is segmented into non-overlapping windows (default M = 7), and self-attention is calculated within each window:
[0141]
[0142] Among them, Q, K, and V are the query, key, and value matrices, and dk is the dimensionality scaling factor.
[0143] Global interaction module: Introduce a cross-window information transfer mechanism, capture long-range dependencies through deformable convolution, and output image features FI ∈ Rdi (di = 1024).
[0144] 2. Text feature extraction (Longformer encoder):
[0145] Use Longformer with a sparse attention mechanism to process long texts, and its attention weight matrix A satisfies:
[0146]
[0147] Where is the window radius, is the preset global attention position. Finally, output the text feature .
[0148] Step S3: Cross-modal contrast learning and feature alignment.
[0149] 1. Feature space mapping:
[0150] Map the image and text features to a unified dimensional space through a linear projection layer:
[0151]
[0152] Where , are learnable parameters, = 512 is the shared dimension.
[0153] 2. InfoNCE contrast loss:
[0154] Define the contrast loss function for the image-text pair:
[0155]
[0156] Where is the similarity score, is the temperature coefficient (dynamic adjustment range is [0.05, 0.2]), and N is the number of negative samples within the batch.
[0157] Step S4: Multi-modal feature fusion and classification.
[0158] 1. Feature fusion:
[0159] Fuse the aligned features using a gated weighting mechanism:
[0160]
[0161] where σ is the Sigmoid function, and ⊙ represents element-wise multiplication, is the gated weight.
[0162] 2. Classifier design:
[0163] The classification head based on Transformer outputs class probabilities:
[0164]
[0165] The cross-entropy loss is , where C is the number of classes (correct wiring, wrong wiring, loose connection, etc.).
[0166] Step S5: Joint training and optimization.
[0167] 1. Joint optimization of the loss function:
[0168] The total loss is the weighted sum of the contrastive loss and the classification loss:
[0169]
[0170] where the weight gradually decays from 0.8 to 0.2 to balance modal alignment and classification accuracy.
[0171] 2. GPU heterogeneous acceleration:
[0172] Image branch: The window attention calculation of MaxViT is implemented through CUDA kernels, and Tensor Core is used to accelerate matrix multiplication (FP16 mixed precision). Text branch: The sparse attention matrix of Longformer is compressed and stored through a mask matrix to reduce video memory occupancy. Parallelization strategy: Feature extraction, contrastive learning, and classification tasks are assigned to a multi-GPU pipeline, and the measured inference speed is increased by 5.3 times (compared with a single V100).
[0173] Step S6: Robustness enhancement and deployment.
[0174] 1. Dynamic temperature coefficient adjustment:
[0175] Adaptively adjust τ according to the difficulty of samples within a batch:
[0176]
[0177] where, = 0.1, α = 0.01 are adjustment coefficients.
[0178] 2. Edge device deployment:
[0179] Through knowledge distillation, MaxViT is compressed into lightweight MobileViT, reducing the number of model parameters by 72%, and achieving real-time detection (>30 FPS) on JetsonAGX Xavier.
[0180] The embodiment of the present application proposes an end-to-end detection framework based on multimodal contrastive learning, innovatively combines the improved Vision Transformer (MaxViT) with the long text encoder (Longformer), and explicitly aligns the semantic space of image features and wiring specification text through cross-modal contrastive learning (InfoNCE loss function); designs a joint optimization strategy to simultaneously optimize modal relevance and classification task specificity, and introduces a dynamic temperature coefficient and an intra-batch negative sampling mechanism to enhance the ability to distinguish subtle wiring errors; at the same time, through lightweight model architecture design and GPU parallel heterogeneous computing acceleration, while ensuring high-precision detection (joint modeling of local details and global topology), the computing delay is significantly reduced, and real-time and robust detection and classification of the wiring status of electric energy metering equipment in complex scenarios are achieved.
[0181] The terms in the above embodiments are explained as follows:
[0182] Multimodal data: refers to perceived information from different sources or types. For example, in this technical solution, multimodal data includes wiring images of electric energy metering equipment and text information describing wiring specifications.
[0183] Contrastive learning: A machine learning method that aims to learn effective feature representations by distinguishing similar and dissimilar samples. In this scheme, contrastive learning is used to bring images and texts closer in semantic space, thereby achieving cross-modal feature alignment.
[0184] InfoNCE loss function: Info Noise Contrastive Estimation, is a commonly used contrastive learning loss function. It learns effective feature representation by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs.
[0185] Vision Transformer (ViT): Vision Transformer, a model that applies the Transformer model to image processing. It divides the image into patches and inputs the patch sequence into the Transformer encoder for feature extraction.
[0186] MaxViT (Maximal Vision Transformer): The Maximal Vision Transformer, an improved Vision Transformer model. It alternately stacks local window attention and global interaction modules to effectively extract local and global features of images, thereby enhancing the model's performance.
[0187] Local Window Attention: A key module in the MaxViT model. It divides the image into non-overlapping windows and calculates self-attention within each window to capture local features.
[0188] Global Interaction Block: Another key module in the MaxViT model. It introduces a cross-window information transfer mechanism and captures long-range dependencies and global information of the image through deformable convolution and other means.
[0189] Longformer: A Transformer model specialized for processing long texts. It adopts a sparse attention mechanism, reducing the computational complexity and enabling it to effectively process long text sequences, such as wiring specification documents.
[0190] Sparse Attention Mechanism: The attention mechanism used in the Longformer model. Different from the global attention of traditional Transformers, sparse attention only focuses on some relevant inputs, thereby reducing the computational amount and improving the efficiency of processing long sequences.
[0191] Pre-trained Word Embeddings: Word vector representations obtained by pre-training on a large-scale text corpus. For example, Word2Vec can map words to a low-dimensional vector space to capture the semantic information of words.
[0192] Feature Space Mapping: The process of projecting feature vectors of different modalities into the same shared feature space. In the embodiments of this application, the image and text features are mapped to a unified dimensional space through a linear projection layer, facilitating subsequent cross-modal contrast learning and fusion.
[0193] Gating Weighting Mechanism: A feature fusion method that dynamically controls the contribution degree of different features by learning gating weights. In the embodiments of this application, the gating weighting mechanism is used to fuse image and text features and adaptively adjust the importance of different modality features.
[0194] Transformer Classification Head: A classifier head based on the Transformer architecture, used to map the fused features to class probabilities.
[0195] Cross-Entropy Loss: A commonly used loss function for classification tasks, used to measure the difference between the class probabilities predicted by the model and the true labels.
[0196] Dynamic Temperature Coefficient: In contrastive learning, the temperature coefficient τ is used to adjust the distribution of similarity scores, affecting the sensitivity of the model to negative samples. The dynamic temperature coefficient can be adaptively adjusted during training to improve model performance.
[0197] Knowledge Distillation: A model compression technique that transfers the knowledge of a complex model (teacher model) to a simple model (student model), enabling the student model to maintain performance close to that of the teacher model while reducing the number of parameters.
[0198] MobileViT: A mobile vision Transformer, a lightweight Vision Transformer model suitable for deployment on mobile and edge devices.
[0199] GPU Heterogeneous Acceleration: Utilize the parallel computing power of the GPU (Graphics Processing Unit) to accelerate the model training and inference process. Heterogeneous acceleration refers to adopting different optimization strategies for different parts of the model to fully utilize the computing resources of the GPU.
[0200] CUDA Kernel: CUDA (Compute Unified Device Architecture) is a parallel computing platform and programming model provided by NVIDIA. A CUDA kernel is a parallel computing function executed on the GPU.
[0201] Tensor Core: A hardware unit in NVIDIA GPUs specifically designed to accelerate matrix multiplication and convolution operations, especially capable of significantly improving computing efficiency in low-precision calculations (such as FP16).
[0202] FP16 Mixed Precision: A way of mixed-precision training that uses both single-precision floating-point numbers (FP32) and half-precision floating-point numbers (FP16) during training. FP16 can reduce the memory footprint and utilize Tensor Core to accelerate calculations, improving the training speed.
[0203] Mask Matrix: In the attention mechanism, the mask matrix is used to mask certain input information to prevent the model from paying attention to positions that should not be focused on. In Longformer, the mask matrix is used to implement sparse attention.
[0204] Multi-GPU Pipeline: Decompose the model training or inference task into multiple stages, and allocate different stages to different GPUs for parallel execution, forming a pipelined computing process to improve the overall processing speed.
[0205] Edge Device Deployment: Deploy the trained model to edge computing devices (such as Jetson AGX Xavier, etc.) for operation to achieve localized real-time detection and processing.
[0206] Real-time Detection: Complete the detection task in an extremely short time to meet the detection application scenarios with real-time requirements. In the embodiments of this application, real-time detection means being able to quickly detect the wiring status of power metering devices and promptly discover abnormalities.
[0207] Robustness: The ability of the model to maintain stable performance when facing various interference factors (such as noise, light changes, etc.).
[0208] Power Metering Device: A device used to measure power consumption, including electricity meters, junction boxes, terminals, etc., which is an important part of the smart grid.
[0209] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0210] Based on the same inventive concept, the embodiments of this application also provide a power metering device wiring specification detection device for implementing the power metering device wiring specification detection method involved above. The solution provided by this device to solve the problem is similar to the solution recorded in the above method. Therefore, the specific limitations in one or more embodiments of the power metering device wiring specification detection device provided below can refer to the limitations on the power metering device wiring specification detection method in the above text, and will not be repeated here.
[0211] In an exemplary embodiment, as Figure 6 shown, a power metering device wiring specification detection device 600 is provided, including: a data acquisition module 601, a feature determination module 602, a feature mapping module 603, and a feature detection module 604, where:
[0212] A data acquisition module 601, configured to acquire a to-be-detected wiring image corresponding to an electric energy metering device, and a wiring specification text corresponding to the to-be-detected wiring image;
[0213] A feature determination module 602, configured to determine an image feature corresponding to the to-be-detected wiring image based on a local window attention module and a global interaction module; determine a text feature corresponding to the wiring specification text based on a sparse attention mechanism; the local window attention module is configured to divide the to-be-detected wiring image into multiple windows, and determine the attention weights of each window; the global interaction module is configured to fuse the image information within the global range through the attention weights of each window to obtain an image feature;
[0214] A feature mapping module 603, configured to map the image feature and the text feature to a unified dimensional space to obtain a first mapping value corresponding to the image feature and a second mapping value corresponding to the text feature;
[0215] A feature detection module 604, configured to fuse the first mapping value and the second mapping value according to a gated weighting mechanism to obtain a fused feature, and determine a wiring detection result corresponding to the fused feature based on a pre-constructed classification head.
[0216] Further, the image feature is a multi-scale image feature. The feature determination module 602 is specifically configured to: divide the to-be-detected wiring image into multiple non-overlapping windows based on the local window attention module and multi-scale configuration information, and determine local attention features corresponding to each window; the local attention feature is determined by weighted summation of the attention weight of the window and the window feature; determine the local attention features corresponding to the multiple non-overlapping windows as single-scale local features to obtain multiple single-scale local features with different scales; the multi-scale configuration information includes the window sizes corresponding to each single-scale local feature; determine the global feature of the to-be-detected wiring image based on the global interaction module; fuse each single-scale local feature with the global feature to obtain a multi-scale image feature.
[0217] Further, the feature determination module 602 is specifically further configured to: determine fixed windows corresponding to each token in the wiring specification text, and determine window text features determined by multiple tokens included in the fixed windows; obtain target tokens with global attention, and determine global text features determined by each target token; fuse each window text feature with the global text feature to obtain a text feature corresponding to the wiring specification text.
[0218] Further, the feature mapping module 603 is specifically configured to: determine a first linear projection layer corresponding to the image feature and a second linear projection layer corresponding to the text feature; the first linear projection layer is used to map the input dimension corresponding to the image feature to a preset dimension; the second linear projection layer is used to map the input dimension corresponding to the text feature to a preset dimension; based on the first linear projection layer, map the image feature to the dimension space corresponding to the preset dimension to obtain a first mapped value corresponding to the image feature; based on the second linear projection layer, map the text feature to the dimension space corresponding to the preset dimension to obtain a second mapped value corresponding to the text feature.
[0219] Further, the feature detection module 604 is specifically configured to: based on the gating weight matrix in the gating weighting mechanism, perform weighted summation on the first mapped value and the second mapped value to obtain a fused feature; input the fused feature into a classification head constructed based on Transformer to obtain probability values of each wiring detection category output by the classification head; based on the probability values of each wiring detection category, determine the wiring detection result corresponding to the fused feature.
[0220] Further, the device further includes a training module, which is specifically configured to: determine a contrast loss function corresponding to the image feature and the text feature, and the formula of the contrast loss function is as follows:
[0221]
[0222]
[0223] Where is the similarity score, is the transpose of the first mapped value corresponding to the image feature, is the second mapped value corresponding to the text feature, is the temperature coefficient, N is the number of negative samples within a batch, is the k-th text negative sample that does not match the first mapped value, is the k-th image negative sample that does not match the second mapped value;
[0224] Determine the cross-entropy loss as the classification loss function; perform weighted summation on the contrast loss function and the classification loss function to obtain a joint loss function; the joint loss function is used to train the model corresponding to the classification head.
[0225] Further, the training module is specifically further configured to: for the contrast loss function, adjust the formula of the temperature coefficient as follows:
[0226]
[0227] Where is the initial temperature value, α is the adjustment coefficient, is the similarity score.
[0228] Each module in the above power metering device wiring specification detection device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0229] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the wiring images to be detected corresponding to the power metering device, as well as the wiring specification texts corresponding to the wiring images to be detected. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for detecting the wiring specification of a power metering device.
[0230] Those skilled in the art can understand that Figure 7 the structure shown in
[0231] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in each of the above method embodiments.
[0232] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in each of the above method embodiments.
[0233] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above-mentioned method embodiments.
[0234] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0235] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0236] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0237] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for detecting wiring specifications of electric energy metering equipment, characterized in that: The method comprises: Acquire a wiring image to be detected corresponding to the electric energy metering device, and a wiring specification text corresponding to the wiring image to be detected; Based on the local window attention module and the global interaction module, the image features corresponding to the wiring image to be detected are determined; based on the sparse attention mechanism, the text features corresponding to the wiring specification text are determined; the local window attention module is used to divide the wiring image to be detected into multiple windows and determine the attention weight of each window; the global interaction module is used to fuse the image information in the global range through the attention weights of each window to obtain the image features; Mapping the image feature and the text feature to a unified dimensional space to obtain a first mapping value corresponding to the image feature and a second mapping value corresponding to the text feature; The first mapping value and the second mapping value are fused according to a gated weighting mechanism to obtain a fused feature, and a wiring detection result corresponding to the fused feature is determined based on a pre-constructed classification head.
2. The method according to claim 1, characterized in that The image feature is a multi-scale image feature; the image feature corresponding to the wiring image to be detected is determined based on the local window attention module and the global interaction module, including: Based on the local window attention module and the multi-scale configuration information, the wiring image to be detected is divided into a plurality of non-overlapping windows, and a local attention feature corresponding to each of the windows is determined; the local attention feature is determined based on the weighted sum of the attention weight of the window and the window feature; Determining that the local attention features corresponding to the multiple non-overlapping windows are single-scale local features, and obtaining multiple single-scale local features of different scales; the multi-scale configuration information includes the window size corresponding to each of the single-scale local features; Determining the global features of the wiring image to be detected based on the global interaction module; Each of the single-scale local features is fused with the global features to obtain a multi-scale image feature.
3. The method according to claim 1, characterized in that: The determining of text features corresponding to the wiring specification text based on the sparse attention mechanism includes: Determine a fixed window corresponding to each word unit in the wiring specification text, and determine a window text feature determined by multiple word units contained in the fixed window; Acquire target word-units with global attention, and determine global text features determined by each of the target word-units; Each of the window text features is fused with the global text feature to obtain text features corresponding to the wiring specification text.
4. The method according to claim 1, characterized in that: The step of mapping the image feature and the text feature to a unified dimensional space to obtain a first mapping value corresponding to the image feature and a second mapping value corresponding to the text feature includes: Determine a first linear projection layer corresponding to the image feature, and determine a second linear projection layer corresponding to the text feature; the first linear projection layer is used to map the input dimension corresponding to the image feature to a preset dimension; the second linear projection layer is used to map the input dimension corresponding to the text feature to the preset dimension; Based on the first linear projection layer, mapping the image feature to the dimensional space corresponding to the preset dimension to obtain a first mapping value corresponding to the image feature; Based on the second linear projection layer, the text feature is mapped to the dimensional space corresponding to the preset dimension to obtain a second mapping value corresponding to the text feature.
5. The method according to claim 1, characterized in that The fusing the first mapping value and the second mapping value according to the gated weighting mechanism to obtain a fused feature, and determining a wiring detection result corresponding to the fused feature based on a pre-built classification head, includes: Based on the gating weight matrix in the gating weighting mechanism, weighted summing the first mapping value and the second mapping value to obtain a fusion feature; Inputting the fused features into a classification head constructed based on Transformer to obtain a probability value of each wiring detection category output by the classification head; Based on the probability values of the wiring detection categories, a wiring detection result corresponding to the fusion feature is determined.
6. The method according to claim 1, characterized in that The method further comprises: Determine the contrast loss function corresponding to the image feature and the text feature, the formula of the contrast loss function is as follows: in, is the similarity score, is the transpose of the first mapping value corresponding to the image feature, is the second mapping value corresponding to the text feature, is the temperature coefficient, N is the number of negative samples in the batch, is the kth text negative sample that does not match the first mapping value, is the kth image negative sample that does not match the second mapping value; Determine the cross entropy loss as the classification loss function; The contrast loss function and the classification loss function are weighted and summed to obtain a joint loss function; the joint loss function is used to train the model corresponding to the classification head.
7. The method according to claim 6, characterized in that The method further comprises: For the contrast loss function, the formula for adjusting the temperature coefficient is as follows: in, is the initial temperature value, α is the adjustment coefficient, is the similarity score.
8. A device for detecting wiring specifications of electric energy metering equipment, characterized in that: The device comprises: A data acquisition module, used to acquire a wiring image to be detected corresponding to the electric energy metering device, and a wiring specification text corresponding to the wiring image to be detected; A feature determination module is used to determine the image features corresponding to the wiring image to be detected based on the local window attention module and the global interaction module; and to determine the text features corresponding to the wiring specification text based on the sparse attention mechanism; the local window attention module is used to divide the wiring image to be detected into multiple windows and determine the attention weight of each window; the global interaction module is used to fuse the image information in the global range through the attention weights of each window to obtain the image features; A feature mapping module, used to map the image feature and the text feature to a unified dimensional space to obtain a first mapping value corresponding to the image feature and a second mapping value corresponding to the text feature; The feature detection module is used to fuse the first mapping value and the second mapping value according to a gated weighting mechanism to obtain a fused feature, and determine a wiring detection result corresponding to the fused feature based on a pre-built classification head.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.