Electric energy metering junction box wiring state identification method and device based on multi-mode large model, computer equipment and storage medium

By adopting a multimodal large model in the electrical energy metering junction box recognition system, combining image and text data for pre-training and iterative training, the problem of low accuracy in wiring state recognition in traditional AI recognition systems is solved, and more efficient and robust wiring state recognition is achieved.

CN120047933APending Publication Date: 2025-05-27CHINA SOUTHERN POWER GRID DIGITAL GRID GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510099370.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When traditional AI recognition systems identify the wiring status of the electrical energy metering junction box, they are affected by factors such as ambient light and shooting angle, resulting in low recognition accuracy.

Method used

The multimodal large model is used to obtain the image and text data of the sample junction box, and the multimodal large model is pre-trained. The image encoder, query transformer and large language model are used to determine the matching, generate and compare the loss values ​​of the images and text, and iterative training is carried out to improve the prediction accuracy of the model.

Benefits of technology

It improves the accuracy of the wiring status of the electrical energy metering junction box, enhances the robustness of the model to image quality fluctuations, reduces the dependence on high-quality labeled data, and achieves more efficient wiring status recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047933A_ABST
    Figure CN120047933A_ABST
Patent Text Reader

Abstract

The invention relates to an electric energy metering junction box wiring state identification method and device based on a multi-mode large model, computer equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring sample image data of a sample electric energy metering junction box, and determining sample text data corresponding to the sample image data; determining a matching loss value, a generation loss value and a comparison loss value according to the multi-modal large model to be trained, the sample image data and the sample text data; according to the matching loss value, the generation loss value and the comparison loss value, training the to-be-trained multi-modal large model to obtain a trained multi-modal large model; and acquiring image data of the to-be-analyzed electric energy metering junction box, and inputting the image data into the trained multi-mode large model to obtain a wiring state identification result of the to-be-analyzed electric energy metering junction box. By adopting the method, the identification accuracy of the wiring state of the electric energy metering junction box can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of power grids, and particularly to a method, device, computer device, computer-readable storage medium, and computer program product for identifying the wiring state of an electric energy metering junction box based on a multimodal large model. Background Art

[0002] In the field of power grids, it is crucial to identify the wiring state of an electric energy metering junction box to ensure the stability of the power system.

[0003] In traditional technologies, when identifying the wiring state of an electric energy metering junction box, it is generally identified through an AI (Artificial Intelligence) identification system. However, due to factors such as ambient light and shooting angle, the obtained images of the junction box may have problems such as blurring, poor lighting, complex background, and improper shooting angle, which are prone to errors and result in a low accuracy in identifying the wiring state of the electric energy metering junction box. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for identifying the wiring state of an electric energy metering junction box based on a multimodal large model, which can improve the accuracy of identifying the wiring state of the electric energy metering junction box.

[0005] In a first aspect, the present application provides a method for identifying the wiring state of an electric energy metering junction box based on a multimodal large model, including:

[0006] Obtaining sample image data of a sample electric energy metering junction box;

[0007] Determining, through a multimodal large model to be trained, the image visual features corresponding to the sample image data, determining the text visual features corresponding to the sample image data according to the image visual features, and determining the sample text data corresponding to the sample image data according to the text visual features;

[0008] Determining a matching loss value, a generation loss value, and a contrast loss value according to the multimodal large model to be trained, the sample image data, and the sample text data;

[0009] Obtaining a target loss value according to the matching loss value, the generation loss value, and the contrast loss value;

[0010] Iteratively training the multimodal large model to be trained according to the target loss value to obtain a trained multimodal large model;

[0011] Obtain the image data of the power metering junction box to be analyzed, and input the image data into the trained multimodal large model to obtain the wiring status recognition result of the power metering junction box to be analyzed.

[0012] In one embodiment, the multimodal large model to be trained includes an image encoder, a query transformer, and a large language model;

[0013] Determining the image visual features corresponding to the sample image data through the multimodal large model to be trained, determining the text visual features corresponding to the sample image data according to the image visual features, and determining the sample text data corresponding to the sample image data according to the text visual features includes:

[0014] Determine the image visual features corresponding to the sample image data through the image encoder;

[0015] Determine the text visual features corresponding to the sample image data through the query transformer according to the image visual features;

[0016] Determine the sample text data corresponding to the sample image data through the large language model according to the text visual features.

[0017] In one embodiment, determining the matching loss value, the generation loss value, and the contrast loss value according to the multimodal large model to be trained, the sample image data, and the sample text data includes:

[0018] Input the sample image data and the sample text data into the multimodal large model to be trained to obtain the matching probability between the sample image data and the sample text data;

[0019] Determine the matching loss value according to the matching probability.

[0020] In one embodiment, determining the matching loss value, the generation loss value, and the contrast loss value according to the multimodal large model to be trained, the sample image data, and the sample text data further includes:

[0021] Perform masking processing on the sample text data to obtain the processed sample text data carrying the masked text data;

[0022] Input the sample image data and the processed sample text data into the multimodal large model to be trained to obtain the prediction probability corresponding to the masked text data in the processed sample text data;

[0023] Determine the generation loss value according to the prediction probability.

[0024] In one embodiment, determining the matching loss value, the generation loss value, and the contrast loss value according to the multi-modal large model to be trained, the sample image data, and the sample text data further includes:

[0025] Determine a first embedding vector corresponding to the sample image data and a second embedding vector corresponding to the sample text data through the multi-modal large model to be trained;

[0026] Determine the similarity between the first embedding vector and the second embedding vector;

[0027] Determine the contrast loss value according to the similarity.

[0028] In one embodiment, obtaining the target loss value according to the matching loss value, the generation loss value, and the contrast loss value includes:

[0029] Obtain a first preset weight corresponding to the matching loss value, a second preset weight corresponding to the generation loss value, and a third preset weight corresponding to the contrast loss value;

[0030] Perform weighted summation processing on the matching loss value, the generation loss value, and the contrast loss value according to the first preset weight, the second preset weight, and the third preset weight to obtain the target loss value.

[0031] In one embodiment, obtaining the sample image data of the sample power metering junction box includes:

[0032] Obtain system image data, on-site image data, and simulation image data associated with the sample power metering junction box;

[0033] Preprocess the system image data, the on-site image data, and the simulation image data to obtain preprocessed system image data, preprocessed on-site image data, and preprocessed simulation image data;

[0034] Use the preprocessed system image data, the preprocessed on-site image data, and the preprocessed simulation image data as the sample image data of the sample power metering junction box.

[0035] In a second aspect, the present application also provides an apparatus for identifying the wiring state of a power metering junction box based on a multi-modal large model, including:

[0036] A data acquisition module for acquiring sample image data of a sample power metering junction box;

[0037] A data determination module, configured to determine the image visual features corresponding to the sample image data through a multi-modal large model to be trained, determine the text visual features corresponding to the sample image data according to the image visual features, and determine the sample text data corresponding to the sample image data according to the text visual features;

[0038] A loss determination module, configured to determine a matching loss value, a generation loss value, and a contrast loss value according to the multi-modal large model to be trained, the sample image data, and the sample text data;

[0039] An objective determination module, configured to obtain an objective loss value according to the matching loss value, the generation loss value, and the contrast loss value;

[0040] A model training module, configured to perform iterative training on the multi-modal large model to be trained according to the objective loss value to obtain a trained multi-modal large model;

[0041] A result determination module, configured to obtain the image data of the power metering junction box to be analyzed and input the image data into the trained multi-modal large model to obtain the wiring status recognition result of the power metering junction box to be analyzed.

[0042] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0043] Obtain the sample image data of the sample power metering junction box;

[0044] Determine the image visual features corresponding to the sample image data through a multi-modal large model to be trained, determine the text visual features corresponding to the sample image data according to the image visual features, and determine the sample text data corresponding to the sample image data according to the text visual features;

[0045] Determine a matching loss value, a generation loss value, and a contrast loss value according to the multi-modal large model to be trained, the sample image data, and the sample text data;

[0046] Obtain an objective loss value according to the matching loss value, the generation loss value, and the contrast loss value;

[0047] Perform iterative training on the multi-modal large model to be trained according to the objective loss value to obtain a trained multi-modal large model;

[0048] Obtain the image data of the power metering junction box to be analyzed, and input the image data into the trained multi-modal large model to obtain the wiring status recognition result of the power metering junction box to be analyzed.

[0049] Fourthly, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0050] Obtain the sample image data of the sample power metering junction box;

[0051] Through the multi-modal large model to be trained, determine the image visual features corresponding to the sample image data. According to the image visual features, determine the text visual features corresponding to the sample image data, and according to the text visual features, determine the sample text data corresponding to the sample image data;

[0052] According to the multi-modal large model to be trained, the sample image data and the sample text data, determine the matching loss value, the generation loss value and the contrast loss value;

[0053] According to the matching loss value, the generation loss value and the contrast loss value, obtain the target loss value;

[0054] According to the target loss value, perform iterative training on the multi-modal large model to be trained to obtain the trained multi-modal large model;

[0055] Obtain the image data of the power metering junction box to be analyzed, and input the image data into the trained multi-modal large model to obtain the wiring status recognition result of the power metering junction box to be analyzed.

[0056] Fifthly, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0057] Obtain the sample image data of the sample power metering junction box;

[0058] Through the multi-modal large model to be trained, determine the image visual features corresponding to the sample image data. According to the image visual features, determine the text visual features corresponding to the sample image data, and according to the text visual features, determine the sample text data corresponding to the sample image data;

[0059] According to the multi-modal large model to be trained, the sample image data and the sample text data, determine the matching loss value, the generation loss value and the contrast loss value;

[0060] Obtaining a target loss value according to the matching loss value, the generation loss value and the comparison loss value;

[0061] Iteratively training the multimodal large model to be trained according to the target loss value to obtain a trained multimodal large model;

[0062] Image data of the electric energy metering junction box to be analyzed is obtained, and the image data is input into the trained multimodal large model to obtain a wiring state recognition result of the electric energy metering junction box to be analyzed.

[0063] The above-mentioned method, device, computer equipment, storage medium and computer program product for identifying the wiring status of an electric energy metering junction box based on a multimodal large model first obtain sample image data of a sample electric energy metering junction box, determine the image visual features corresponding to the sample image data through the multimodal large model to be trained, determine the text visual features corresponding to the sample image data based on the image visual features, and determine the sample text data corresponding to the sample image data based on the text visual features, and then determine the matching loss value, generation loss value and contrast loss value based on the multimodal large model to be trained, the sample image data and the sample text data, then obtain the target loss value based on the matching loss value, the generation loss value and the contrast loss value, and then, iteratively train the multimodal large model to be trained based on the target loss value to obtain a trained multimodal large model, and finally, obtain the image data of the electric energy metering junction box to be analyzed, and input the image data into the trained multimodal large model to obtain the wiring status recognition result of the electric energy metering junction box to be analyzed. In this way, when identifying the wiring status of the electric energy metering junction box, the multimodal large model is pre-trained by using the sample image data and sample text data of the sample electric energy metering junction box, so that in practical applications, after obtaining the image data of the electric energy metering junction box to be analyzed, the wiring status recognition result of the electric energy metering junction box to be analyzed can be predicted; secondly, the multimodal large model receives new data in each round of iterative training, so as to improve and optimize the internal model, so as to be able to predict more effectively, which is conducive to improving the prediction accuracy of the multimodal large model, and then improving the recognition accuracy of the wiring status of the electric energy metering junction box; moreover, the whole process does not require human intervention, avoiding the defect of subjective factors and easy errors in the method of artificial visual detection, resulting in low recognition accuracy of the wiring status of the electric energy metering junction box, and further improving the recognition accuracy of the wiring status of the electric energy metering junction box. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] To more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0065] Figure 1 It is a schematic flowchart of a method for identifying the wiring state of an electric energy metering junction box based on a multimodal large model in one embodiment;

[0066] Figure 2 It is a schematic flowchart of a method for identifying the wiring state of an electric energy metering junction box based on a multimodal large model in another embodiment;

[0067] Figure 3 It is a schematic flowchart of a method for identifying abnormal wiring methods of an electric energy metering junction box based on a multimodal large model in one embodiment;

[0068] Figure 4 It is a schematic diagram of the wiring specification of a three-phase three-wire junction box in one embodiment;

[0069] Figure 5 It is a schematic diagram of the wiring specification of a three-phase four-wire junction box in one embodiment;

[0070] Figure 6 It is a schematic diagram of the structure of the BLIP-2 model in one embodiment;

[0071] Figure 7 It is a schematic diagram of the structure of the BLIP-2 model in another embodiment;

[0072] Figure 8 It is a schematic diagram of the structure of the BLIP-2 model in yet another embodiment;

[0073] Figure 9 It is a structural block diagram of a device for identifying the wiring state of an electric energy metering junction box based on a multimodal large model in one embodiment;

[0074] Figure 10 It is an internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0075] In order to make the purpose, technical solutions and advantages of the present application more clear and understandable, the following further details the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0076] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0077] In an exemplary embodiment, as Figure 1 shown, a method for identifying the wiring status of an electric energy metering junction box based on a multi-modal large model is provided. In this embodiment, an example is given where this method is applied to a server; it can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, and tablet computers; the server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:

[0078] Step S101, obtain sample image data of a sample electric energy metering junction box.

[0079] Among them, the sample electric energy metering junction box refers to an electric energy metering junction box used for model training.

[0080] Among them, the electric energy metering junction box refers to an electrical device used to connect and distribute the secondary circuit wires of current transformers and voltage transformers in an electric energy metering device, such as a three-phase three-wire junction box, a three-phase four-wire junction box, etc.

[0081] Among them, the sample image data refers to image data associated with the sample electric energy metering junction box.

[0082] Exemplarily, in response to a model training instruction for a multi-modal large model to be trained, the server obtains system image data, on-site image data, and simulation image data associated with the sample electric energy metering junction box, all of which are used as the sample image data of the sample electric energy metering junction box.

[0083] Step S102, through the multi-modal large model to be trained, determine the image visual features corresponding to the sample image data, determine the text visual features corresponding to the sample image data according to the image visual features, and determine the sample text data corresponding to the sample image data according to the text visual features.

[0084] Among them, the multimodal large model refers to a network model that can utilize the image data of the electric energy metering junction box to obtain the recognition result of the wiring state of the electric energy metering junction box to be analyzed, such as the BLIP-2 (Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models) model.

[0085] Among them, the image visual features are used to represent the relevant attribute information reflecting the characteristics of the image content of the sample image data at the visual level, such as color features, texture features, etc.

[0086] Among them, the text visual features are the relevant attribute information reflecting the characteristics of the text content of the sample text data at the visual level, such as font type, font size, etc.

[0087] Among them, the sample text data is used to characterize the text data corresponding to the wiring state of the sample electric energy metering junction box, including text data such as junction box description, wiring state judgment conclusion, wiring color judgment conclusion, detailed description of wiring state and wiring color, etc.

[0088] Exemplarily, the server inputs the sample image data of the sample electric energy metering junction box into the multimodal large model to be trained, and determines the image visual features corresponding to the sample image data through the multimodal large model to be trained; then, the server determines the text visual features corresponding to the image visual features through the multimodal large model to be trained as the text visual features corresponding to the sample image data; then, the server determines the text data corresponding to the text visual features through the multimodal large model to be trained according to the text visual features as the sample text data corresponding to the sample image data.

[0089] Step S103, determine the matching loss value, generation loss value, and contrast loss value according to the multimodal large model to be trained, the sample image data, and the sample text data.

[0090] Among them, the matching loss value is a quantization index used to evaluate whether an image and text belong to the same instance.

[0091] Among them, the generation loss value is used to evaluate the model's ability to generate text.

[0092] Among them, the contrast loss value is a quantization index used to measure the degree of difference between an image and text.

[0093] Exemplarily, the server obtains the multi-modal large model to be trained and performs initialization processing on the multi-modal large model to be trained to obtain the processed multi-modal large model; then, the server determines the matching loss value, the generation loss value, and the contrast loss value according to the multi-modal large model, the sample image data, and the sample text data.

[0094] Step S104, obtain the target loss value according to the matching loss value, the generation loss value, and the contrast loss value.

[0095] Among them, the target loss value is used to represent the comprehensive loss value determined according to the multi-modal large model to be trained, the sample image data, and the sample text data.

[0096] Exemplarily, the server constructs the corresponding relationship between the matching loss value, the generation loss value, the contrast loss value, and the target loss value as the target corresponding relationship; then, the server queries the target corresponding relationship according to the matching loss value, the generation loss value, and the contrast loss value to obtain the target loss value corresponding to the matching loss value, the generation loss value, and the contrast loss value.

[0097] Step S105, perform iterative training on the multi-modal large model to be trained according to the target loss value to obtain the trained multi-modal large model.

[0098] Exemplarily, the server adjusts the model parameters of the multi-modal large model to be trained according to the target loss value; then, the server retrains the multi-modal large model with the adjusted model parameters until the target loss value obtained by the trained multi-modal large model is less than the loss value threshold, then stops training, and uses the trained multi-modal large model as the trained multi-modal large model.

[0099] Step S106, obtain the image data of the power metering junction box to be analyzed, and input the image data into the trained multi-modal large model to obtain the wiring state recognition result of the power metering junction box to be analyzed.

[0100] Among them, the power metering junction box to be analyzed refers to the power metering junction box that needs to be recognized for the wiring state.

[0101] Among them, the wiring state recognition result is used to represent the wiring state of the power metering junction box to be analyzed.

[0102] Exemplarily, the server responds to the wiring state recognition instruction of the power metering junction box to be analyzed, obtains the image data of the power metering junction box to be analyzed from the database; then, the server inputs the image data into the trained multi-modal large model to obtain the text data corresponding to the image data as the wiring state recognition result of the power metering junction box to be analyzed.

[0103] In the above-mentioned method for identifying the wiring state of an electric energy metering junction box based on a multimodal large model, sample image data of a sample electric energy metering junction box is first obtained, and the image visual features corresponding to the sample image data are determined through the multimodal large model to be trained. According to the image visual features, the text visual features corresponding to the sample image data are determined, and according to the text visual features, the sample text data corresponding to the sample image data is determined. Then, according to the multimodal large model to be trained, the sample image data and the sample text data, the matching loss value, the generation loss value and the contrast loss value are determined. Then, according to the matching loss value, the generation loss value and the contrast loss value, the target loss value is obtained. Then, according to the target loss value, the multimodal large model to be trained is iteratively trained to obtain a trained multimodal large model. Finally, the image data of the electric energy metering junction box to be analyzed is obtained, and the image data is input into the trained multimodal large model to obtain the wiring state recognition result of the electric energy metering junction box to be analyzed. In this way, when identifying the wiring status of the electric energy metering junction box, the multimodal large model is pre-trained by using the sample image data and sample text data of the sample electric energy metering junction box, so that in practical applications, after obtaining the image data of the electric energy metering junction box to be analyzed, the wiring status recognition result of the electric energy metering junction box to be analyzed can be predicted; secondly, the multimodal large model receives new data in each round of iterative training, so as to improve and optimize the internal model, so as to be able to predict more effectively, which is conducive to improving the prediction accuracy of the multimodal large model, and then improving the recognition accuracy of the wiring status of the electric energy metering junction box; moreover, the whole process does not require human intervention, avoiding the defect of subjective factors and easy errors in the method of artificial visual detection, resulting in low recognition accuracy of the wiring status of the electric energy metering junction box, and further improving the recognition accuracy of the wiring status of the electric energy metering junction box.

[0104] In an exemplary embodiment, the multimodal large model to be trained includes an image encoder, a query transformer, and a large language model.

[0105] Then, the above step S102, through the multimodal large model to be trained, determines the image visual features corresponding to the sample image data, determines the text visual features corresponding to the sample image data based on the image visual features, and determines the sample text data corresponding to the sample image data based on the text visual features, specifically includes the following contents: determining the image visual features corresponding to the sample image data through the image encoder; determining the text visual features corresponding to the sample image data through the query transformer based on the image visual features; determining the sample text data corresponding to the sample image data based on the text visual features through the large language model.

[0106] The image encoder is also called Image Encoder.

[0107] Among them, the query transformer is also called Q-Former (Query Former).

[0108] Among them, the large language model is also called LLM (Large Language Model).

[0109] Exemplarily, the server determines the image visual features corresponding to the sample image data through the image encoder; then, the server determines the text visual features corresponding to the image visual features, as the text visual features corresponding to the sample image data, through the query transformer according to the image visual features and the preset text (such as a question sentence); then, the server determines the text data corresponding to the text visual features, as the sample text data corresponding to the sample image data, through the large language model according to the text visual features.

[0110] In this embodiment, by first extracting the image visual features, then determining the text visual features based on this, and finally generating the sample text data, the visual information contained in the image can be presented in the form of text, so that information of different modalities can be mutually integrated and supplemented, facilitating subsequent data processing.

[0111] In an exemplary embodiment, in step S103 above, according to the multi-modal large model to be trained, the sample image data, and the sample text data, a matching loss value, a generation loss value, and a contrast loss value are determined, which specifically include the following: inputting the sample image data and the sample text data into the multi-modal large model to be trained to obtain the matching probability between the sample image data and the sample text data; determining the matching loss value according to the matching probability.

[0112] Among them, the matching probability is used to represent the possibility that the sample image data and the sample text data belong to the same instance.

[0113] Exemplarily, the server combines the sample image data and the sample text data to obtain multiple groups of image-text pairs; for example, the server combines the sample image data 1 and the sample text data 1, combines the sample image data 1 and the sample text data 2, etc., to obtain multiple groups of image-text pairs; then, the server inputs the multiple groups of image-text pairs into the multi-modal large model to be trained, and through the multi-modal large model, obtains the matching probability between the sample image data and the sample text data in each group of image-text pairs; then, the server determines the binary label of each group of image-text pairs; for example, if the image-text pair obtained by combining the sample image data 1 and the sample text data 1 is a matching image-text pair, the binary label of this image-text pair is set to 1; if the image-text pair obtained by combining the sample image data 1 and the sample text data 2 is a non-matching image-text pair, the binary label of this image-text pair is set to 0; then, the server queries the corresponding relationship between the matching probability, the number, the binary label and the matching loss value according to the matching probability between the sample image data and the sample text data in each group of image-text pairs, the sample number of the image-text pair, and the binary label of each group of image-text pairs, and obtains the matching loss value.

[0114] For example, the server can obtain the matching loss value through the following formula:

[0115]

[0116] where L ML refers to the matching loss value; I i refers to the sample image data; T i refers to the sample text data corresponding to the sample image data; N refers to the sample number of the image-text pair (I i , T i ); y i is the binary label, and a value of 1 indicates that (I i , T i ) is a matching image-text pair, and a value of 0 indicates a non-matching image-text pair; p match refers to the matching probability.

[0117] In this embodiment, by determining the matching loss value according to the matching probability between the sample image data and the sample text data, the matching ability of the multi-modal large model for images and texts can be monitored, thereby improving the reliability of the recognition result of the multi-modal large model.

[0118] In an exemplary embodiment, in step S103 above, according to the multi-modal large model to be trained, the sample image data, and the sample text data, a matching loss value, a generation loss value, and a contrast loss value are determined. Specifically, it includes the following: performing masking processing on the sample text data to obtain the processed sample text data carrying masked text data; inputting the sample image data and the processed sample text data into the multi-modal large model to be trained to obtain the prediction probability corresponding to the masked text data in the processed sample text data; and determining the generation loss value according to the prediction probability.

[0119] Among them, the masking processing refers to the process of replacing a part of the text data in the sample text data with special symbols (such as "[MASK]").

[0120] Among them, the masked text data is used to represent the masked words replaced by special symbols.

[0121] Among them, the prediction probability is used to represent the possibility determined by the multi-modal large model corresponding to the masked text data.

[0122] Exemplarily, the server performs word segmentation processing on the sample text data to obtain multiple word segments corresponding to the sample text data; then, the server determines the importance degree corresponding to each word segment, and filters out the word segments with the importance degree greater than the preset importance degree from each word segment as the target word segments; then, the server uses the preset special symbols to replace the target word segments in the sample text data to obtain the masked text data; then, the server inputs the sample image data and the processed sample text data carrying the masked text data into the multi-modal large model to be trained, and through the multi-modal large model, obtains multiple predicted text data; then, the server determines the predicted text data that matches the masked text data from each predicted text data as the target predicted text data; then, the server obtains the prediction probability corresponding to the target predicted text data, and queries the corresponding relationship between the prediction probability and the generation loss value according to the prediction probability to obtain the generation loss value.

[0123] For example, the server can obtain the generation loss value through the following formula:

[0124] L GL =-∑ t∈masked positions log p(ω t |T′), Equation (2)

[0125] Among them, L GL refers to the generation loss value; t is used to represent the position of the masked text data in the processed sample text data; p(ω t |T′) refers to the prediction probability of the target predicted text data ω t corresponding to the masked text data T′.

[0126] In this embodiment, by performing masking processing on the sample text data, a certain degree of uncertainty is introduced into the model, so that the multi-modal large model cannot solely rely on specific sample text data during the training process, but also needs to learn more general patterns and features, which is beneficial to improving the adaptability and generalization ability of the multi-modal large model to different inputs.

[0127] In an exemplary embodiment, in step S103 above, according to the multi-modal large model to be trained, the sample image data, and the sample text data, a matching loss value, a generation loss value, and a contrast loss value are determined, which specifically includes the following: Through the multi-modal large model to be trained, a first embedding vector corresponding to the sample image data and a second embedding vector corresponding to the sample text data are determined; the similarity between the first embedding vector and the second embedding vector is determined; and according to the similarity, the contrast loss value is determined.

[0128] Among them, the first embedding vector refers to the embedding vector corresponding to the sample image data.

[0129] Among them, the second embedding vector refers to the embedding vector corresponding to the sample text data.

[0130] Among them, the similarity is used to characterize the similarity degree between the first embedding vector and the second embedding vector, and it can be the cosine similarity.

[0131] Exemplarily, the server performs feature extraction processing on the sample image data through the multi-modal large model to be trained to obtain the embedding vector corresponding to the sample image data as the first embedding vector, and performs feature extraction processing on the sample text data to obtain the embedding vector corresponding to the sample text data as the second embedding vector; then, the server determines the similarity between the first embedding vector and the second embedding vector as the similarity score between the first embedding vector and the second embedding vector; and then, the server queries the corresponding relationship between the similarity score and the contrast loss value according to the similarity score to obtain the contrast loss value.

[0132] Illustratively, the server can obtain the contrast loss value through the following formula:

[0133]

[0134] Among them, L CL refers to the contrast loss value; N refers to the combination number between the first embedding vector and the second embedding vector; T j refers to any one of the sample text data; s(I i , T i ) is used to represent the image-text pair (I i , T i) similarity score; τ refers to the temperature parameter used to adjust the similarity score distribution, generally taking values between 0.07 and 0.1.

[0135] In this embodiment, by determining the similarity between the first embedding vector and the second embedding vector, the correlation degree between the image and the text can be quantified; moreover, by comparing the similarity of the embedding vectors of different samples, the model can learn the differences and similarities between different wiring states, which is beneficial to improving the prediction accuracy of the multimodal large model.

[0136] In an exemplary embodiment, in step S104 above, according to the matching loss value, the generation loss value, and the contrast loss value, the target loss value is obtained, which specifically includes the following: obtaining the first preset weight corresponding to the matching loss value, the second preset weight corresponding to the generation loss value, and the third preset weight corresponding to the contrast loss value; according to the first preset weight, the second preset weight, and the third preset weight, performing a weighted sum processing on the matching loss value, the generation loss value, and the contrast loss value to obtain the target loss value.

[0137] Among them, the first preset weight refers to the preset weight corresponding to the matching loss value, such as 0.5.

[0138] Among them, the preset weight refers to the preset weight coefficient. It should be noted that the preset weight can be determined according to the situation.

[0139] Among them, the second preset weight refers to the preset weight corresponding to the generation loss value, such as 0.3.

[0140] Among them, the third preset weight refers to the preset weight corresponding to the contrast loss value, such as 0.2.

[0141] Exemplarily, the server determines the initial weight corresponding to the matching loss value according to the importance of the matching loss value, determines the initial weight corresponding to the generation loss value according to the importance of the generation loss value, and determines the initial weight corresponding to the contrast loss value according to the importance of the contrast loss value; then, the server performs a normalization process on the initial weight corresponding to the matching loss value, the initial weight corresponding to the generation loss value, and the initial weight corresponding to the contrast loss value to obtain the first preset weight corresponding to the matching loss value, the second preset weight corresponding to the generation loss value, and the third preset weight corresponding to the contrast loss value; then, the server performs a weighted sum processing on the matching loss value, the generation loss value, and the contrast loss value according to the first preset weight, the second preset weight, and the third preset weight to obtain the loss value after the weighted sum processing as the target loss value.

[0142] In this embodiment, by reasonably setting preset weights, weighted summation processing is performed on three different loss values to obtain a target loss value, thereby balancing the relationship between different loss values, avoiding a certain loss value from overly dominating the training process, and further improving the stability of the multi-modal large model training.

[0143] In an exemplary embodiment, the above step S101 of obtaining the sample image data of the sample electric energy metering junction box specifically includes the following contents: obtaining the system image data, on-site image data, and simulation image data associated with the sample electric energy metering junction box; performing preprocessing on the system image data, on-site image data, and simulation image data to obtain the preprocessed system image data, preprocessed on-site image data, and preprocessed simulation image data; and using the preprocessed system image data, preprocessed on-site image data, and preprocessed simulation image data as the sample image data of the sample electric energy metering junction box.

[0144] Among them, the system image data is the image data containing the sample electric energy metering junction box saved in the existing system, including but not limited to the image data during the installation, maintenance, and inspection of the metering device. It should be noted that the existing system refers to the system that has been deployed and is running, such as the marketing system or the mobile operation APP (application).

[0145] Among them, the on-site image data refers to the image data obtained by photographing the sample electric energy metering junction box in the actual application environment, including the image data corresponding to various wiring methods, wiring environments, and possible abnormal states.

[0146] Among them, the simulation image data is the image data related to the electric energy metering junction box generated in the laboratory environment through a specific simulation environment (such as lighting conditions, background, and the placement angle of the junction box).

[0147] Among them, the preprocessed system image data refers to the system image data after preprocessing.

[0148] Among them, the preprocessed on-site image data refers to the on-site image data after preprocessing.

[0149] Among them, the preprocessed simulation image data refers to the simulation image data after preprocessing.

[0150] Exemplarily, the server obtains system image data, on-site image data, and simulation image data associated with the sample power metering junction box from the database; then, the server preprocesses the system image data, on-site image data, and simulation image data to obtain preprocessed system image data, preprocessed on-site image data, and preprocessed simulation image data; for example, the server screens the system image data, on-site image data, and simulation image data according to the data acquisition specification to obtain screened system image data, screened on-site image data, and screened simulation image data, and denoises the screened system image data, screened on-site image data, and screened simulation image data to obtain preprocessed system image data, preprocessed on-site image data, and preprocessed simulation image data; then, the server combines the preprocessed system image data, preprocessed on-site image data, and preprocessed simulation image data to obtain the sample image data of the sample power metering junction box.

[0151] In this embodiment, obtaining the image data of the sample power metering junction box through multiple channels is beneficial to improving the diversity of the sample image data and providing more comprehensive learning materials for the multimodal large model; moreover, preprocessing the image data further improves the quality of the sample image data and provides a good data basis for the subsequent training process.

[0152] In an exemplary embodiment, as Figure 2 shown, another method for identifying the wiring state of a power metering junction box based on a multimodal large model is provided. Taking this method applied to a server as an example, the method includes the following steps:

[0153] Step S201, obtain system image data, on-site image data, and simulation image data associated with the sample power metering junction box.

[0154] Step S202, preprocess the system image data, on-site image data, and simulation image data to obtain preprocessed system image data, preprocessed on-site image data, and preprocessed simulation image data.

[0155] Step S203, use the preprocessed system image data, preprocessed on-site image data, and preprocessed simulation image data as the sample image data of the sample power metering junction box.

[0156] Step S204, through the multimodal large model to be trained, determine the image visual features corresponding to the sample image data, according to the image visual features, determine the text visual features corresponding to the sample image data, and according to the text visual features, determine the sample text data corresponding to the sample image data.

[0157] Step S205: Input the sample image data and the sample text data into the multi-modal large model to be trained, and obtain the matching probability between the sample image data and the sample text data; determine the matching loss value according to the matching probability.

[0158] Step S206: Perform masking processing on the sample text data to obtain the processed sample text data carrying the masked text data; input the sample image data and the processed sample text data into the multi-modal large model to be trained, and obtain the prediction probability corresponding to the masked text data in the processed sample text data; determine the generation loss value according to the prediction probability.

[0159] Step S207: Determine the first embedding vector corresponding to the sample image data and the second embedding vector corresponding to the sample text data through the multi-modal large model to be trained; determine the similarity between the first embedding vector and the second embedding vector; determine the contrast loss value according to the similarity.

[0160] Step S208: Obtain the first preset weight corresponding to the matching loss value, the second preset weight corresponding to the generation loss value, and the third preset weight corresponding to the contrast loss value; perform weighted summation processing on the matching loss value, the generation loss value, and the contrast loss value according to the first preset weight, the second preset weight, and the third preset weight to obtain the target loss value.

[0161] Step S209: Iteratively train the multi-modal large model to be trained according to the target loss value to obtain the trained multi-modal large model.

[0162] Step S210: Obtain the image data of the power metering junction box to be analyzed, and input the image data into the trained multi-modal large model to obtain the wiring state recognition result of the power metering junction box to be analyzed.

[0163] In the above method for identifying the wiring state of an electric energy metering junction box based on a multi-modal large model, when identifying the wiring state of the electric energy metering junction box, the multi-modal large model is pre-trained by using the sample image data and sample text data of the sample electric energy metering junction box. In practical applications, after obtaining the image data of the electric energy metering junction box to be analyzed, it is only necessary to predict the identification result of the wiring state of the electric energy metering junction box to be analyzed. Secondly, the multi-modal large model receives new data in each round of iterative training, so as to improve and optimize the model internally, making it possible to predict more effectively, which is conducive to improving the prediction accuracy of the multi-modal large model, and further improving the identification accuracy of the wiring state of the electric energy metering junction box. Moreover, the entire process does not require manual intervention, avoiding the defect that the subjective factors in the method of manual visual inspection are likely to cause errors and result in a low identification accuracy of the wiring state of the electric energy metering junction box, and further improving the identification accuracy of the wiring state of the electric energy metering junction box.

[0164] In an exemplary embodiment, in order to more clearly illustrate the method for identifying the wiring state of an electric energy metering junction box based on a multi-modal large model provided by the embodiments of the present application, the following uses a specific embodiment to specifically describe the method for identifying the wiring state of an electric energy metering junction box based on a multi-modal large model. In one embodiment, as Figure 3 shown, the present application also provides an algorithm for identifying abnormal wiring methods of an electric energy metering junction box based on a multi-modal large model. When identifying the wiring state of the electric energy metering junction box, first obtain the sample image data of the sample electric energy metering junction box, determine the image visual features corresponding to the sample image data through the multi-modal large model to be trained, determine the text visual features corresponding to the sample image data according to the image visual features, and determine the sample text data corresponding to the sample image data according to the text visual features. Then, according to the multi-modal large model to be trained, the sample image data and the sample text data, determine the matching loss value, the generation loss value and the contrast loss value. Next, according to the matching loss value, the generation loss value and the contrast loss value, obtain the target loss value. Then, according to the target loss value, perform iterative training on the multi-modal large model to be trained to obtain a trained multi-modal large model. Finally, obtain the image data of the electric energy metering junction box to be analyzed, and input the image data into the trained multi-modal large model to obtain the identification result of the wiring state of the electric energy metering junction box to be analyzed. The specific content is as follows:

[0165] 1. Data collection.

[0166] 1.1 Multi-source collection:

[0167] (1) Integration of existing system data:

[0168] The existing systems refer to the systems that have been deployed and are running, such as marketing systems or mobile operation APPs. They have accumulated some data on the wiring boxes of metering devices during their daily operations, including but not limited to image data during the installation, repair, and inspection of metering devices. Integrating these data can provide the image of the wiring status in the real scenario for the project.

[0169] (2) On-site shooting:

[0170] On-site shooting refers to shooting the wiring box in the actual application environment. On-site shooting can provide the most real and direct usage situation of the wiring box, including various wiring methods, wiring environments, and possible abnormal states. When conducting on-site shooting, special attention should be paid to the diversity of the shooting environment to ensure that the collected images can cover different lighting conditions, background environments, and types of wiring boxes.

[0171] (3) Laboratory simulation:

[0172] Laboratory simulation is a data collection method carried out in a controlled environment. By setting up a simulated environment in the laboratory, factors such as the lighting conditions, background, and placement angle of the wiring box for shooting can be precisely controlled to produce high-quality data samples. The main advantage of laboratory simulation is that it can systematically construct wiring box images including various abnormal situations, such as deliberately creating common faults like wiring errors and loose connections. This method is suitable for generating abnormal wiring situations that are difficult to obtain in the on-site environment, thereby enhancing the diversity and representativeness of abnormal samples in the dataset.

[0173] 1.2. Data collection review:

[0174] To ensure the quality of the dataset and its effectiveness for subsequent model training, a data screening link needs to be set up to screen data from different sources according to the data collection specifications, further improving the quality of the dataset, ensuring the diversity and representativeness of the data, and providing a solid data foundation for data annotation and subsequent model evaluation.

[0175] 2. Data annotation.

[0176] 2.1. Graphic and text annotation:

[0177] For each picture of the electrical wiring box, judge whether there is an abnormal wiring state in the wiring box shown in the picture, and give a description in the following format:

[0178] {Description of the wiring box:}

[0179] {Judgment conclusion of the wiring state:}

[0180] {Judgment conclusion of the wiring color: Normal wiring color}

[0181] {Detailed description of wiring status and wiring color:}

[0182] Among them, the wiring specification status of three-phase three-wire and three-phase four-wire junction boxes is as follows:

[0183] Wiring specification of three-phase three-wire junction box:

[0184] As Figure 4 shown in the three-phase three-wire junction box, where the junction box structure description: the upper part as a whole is the outgoing terminal position, and the lower part is the incoming terminal position; area 1 is the A-phase voltage area, area 2 is the A-phase current area, area 3 is the B-phase area, area 4 is the C-phase current area, and area 5 is the C-phase voltage area.

[0185] Wiring color description: A-phase is yellow; B-phase is green; C-phase is red.

[0186] Wiring status description:

[0187] (1) Correct wiring status of A-phase voltage area: There is one wire connected to the terminal position at both the incoming terminal position and the outgoing terminal position;

[0188] (2) Correct wiring status of A-phase current area: There are two wires connected at both the incoming and outgoing terminal positions. Among them, the wiring at the incoming terminal position is connected to the second and third terminal positions, and the wiring at the outgoing terminal position is connected to the first and third terminal positions;

[0189] (3) Correct wiring status of B-phase area: There is one wire connected to the terminal position at both the incoming terminal position and the outgoing terminal position;

[0190] (4) Correct wiring status of C-phase current area: There are two wires connected at both the incoming and outgoing terminal positions. Among them, the wiring at the incoming terminal position is connected to the second and third terminal positions, and the wiring at the outgoing terminal position is connected to the first and third terminal positions;

[0191] (5) Correct wiring status of C-phase voltage area: There is one wire connected to the terminal position at both the incoming terminal position and the outgoing terminal position.

[0192] Wiring specification of three-phase four-wire junction box:

[0193] As Figure 5 shown in the three-phase four-wire junction box, where the junction box structure description: the upper part as a whole is the outgoing terminal position, and the lower part is the incoming terminal position; area 1 is the A-phase voltage area, area 2 is the A-phase current area, area 3 is the B-phase voltage area, area 4 is the B-phase current area, area 5 is the C-phase voltage area, area 6 is the C-phase current area, and area 7 is the neutral line area.

[0194] Wiring color description: A-phase is yellow; B-phase is green; C-phase is red; the neutral line is black or blue.

[0195] Description of wiring status:

[0196] (1) Correct wiring status in the A-phase voltage area: One wire is connected to each of the incoming terminal position and the outgoing terminal position on the terminal block;

[0197] (2) Correct wiring status in the A-phase current area: Two wires are connected to both the incoming and outgoing terminal positions. Among them, the wiring at the incoming terminal position is connected to the second and third terminal positions, and the wiring at the outgoing terminal position is connected to the first and third terminal positions;

[0198] (3) Correct wiring status in the B-phase voltage area: One wire is connected to each of the incoming terminal position and the outgoing terminal position on the terminal block;

[0199] (4) Correct wiring status in the B-phase current area: Two wires are connected to both the incoming and outgoing terminal positions. Among them, the wiring at the incoming terminal position is connected to the second and third terminal positions, and the wiring at the outgoing terminal position is connected to the first and third terminal positions;

[0200] (5) Correct wiring status in the C-phase voltage area: One wire is connected to each of the incoming terminal position and the outgoing terminal position on the terminal block;

[0201] (6) Correct wiring status in the C-phase current area: Two wires are connected to both the incoming and outgoing terminal positions. Among them, the wiring at the incoming terminal position is connected to the second and third terminal positions, and the wiring at the outgoing terminal position is connected to the first and third terminal positions;

[0202] (7) Correct wiring status in the neutral line area: One wire is connected to each of the incoming terminal position and the outgoing terminal position on the terminal block, and this wire is the neutral line.

[0203] Examples are as follows:

[0204] Example 1:

[0205] {Description of the junction box: The junction box appears in the figure. The type of the junction box is a three-phase four-wire one-in-one-out junction box, and the wiring method is a three-phase four-wire wiring method}

[0206] {Conclusion of wiring status judgment: Normal wiring status}

[0207] {Conclusion of wiring color judgment: Normal wiring color}

[0208] {Detailed description of wiring status and wiring color:

[0209] One wire is connected to the incoming terminal position in the A-phase voltage area, and the wiring status is correct; the wiring color is yellow, and the wiring color is correct.

[0210] One wire is connected to the outgoing terminal position in the A-phase voltage area, and the wiring status is correct; the wiring color is yellow, and the wiring color is correct.

[0211] There are two wires connected to the second and third incoming terminal positions in the A-phase current area, and the wiring status is correct; the wiring colors are both yellow, and the wiring colors are correct.

[0212] There are two wires connected to the first and third outgoing terminal positions in the A-phase current area, and the wiring status is correct; the wiring colors are both yellow, and the wiring colors are correct.

[0213] There is one wire connected to the incoming terminal position in the B-phase voltage area, and the wiring status is correct; the wiring color is green, and the wiring color is correct.

[0214] There is one wire connected to the outgoing terminal position in the B-phase voltage area, and the wiring status is correct; the wiring color is green, and the wiring color is correct.

[0215] There are two wires connected to the second and third incoming terminal positions in the B-phase current area, and the wiring status is correct; the wiring colors are both green, and the wiring colors are correct.

[0216] There are two wires connected to the first and third outgoing terminal positions in the B-phase current area, and the wiring status is correct; the wiring colors are both green, and the wiring colors are correct.

[0217] There is one wire connected to the incoming terminal position in the C-phase voltage area, and the wiring status is correct; the wiring color is red, and the wiring color is correct.

[0218] There is one wire connected to the outgoing terminal position in the C-phase voltage area, and the wiring status is correct; the wiring color is red, and the wiring color is correct.

[0219] There are two wires connected to the second and third incoming terminal positions in the C-phase current area, and the wiring status is correct; the wiring colors are both red, and the wiring colors are correct.

[0220] There are two wires connected to the first and third outgoing terminal positions in the C-phase current area, and the wiring status is correct; the wiring colors are both red, and the wiring colors are correct.

[0221] There is one wire connected to the incoming terminal position in the neutral wire area, and the wiring status is correct; the wiring color is blue, and the wiring color is correct.

[0222] There is one wire connected to the outgoing terminal position in the neutral wire area, and the wiring status is correct; the wiring color is blue, and the wiring color is correct.

[0223] }

[0224] Example 2:

[0225] {Description of the junction box: The junction box appears in the figure. The type of the junction box is a three-phase four-wire one-in-one-out junction box, and the wiring method is a three-phase four-wire wiring method}

[0226] {Conclusion of wiring status judgment: Abnormal wiring status}

[0227] {Conclusion of wiring color judgment: Normal wiring color}

[0228] {Detailed description of wiring status and wiring color:

[0229] There is a wire connected to the incoming terminal of the A-phase voltage area, and the wiring status is correct; the wiring color is yellow, and the wiring color is correct.

[0230] There is a wire connected to the outgoing terminal of the A-phase voltage area, and the wiring status is correct; the wiring color is yellow, and the wiring color is correct.

[0231] There are two wires connected to the first and third incoming terminals of the A-phase current area. The wiring status is incorrect. It should be connected to the second and third incoming terminals of the A-phase current area; the wiring colors are both yellow, and the wiring colors are correct.

[0232] There are two wires connected to the first and second outgoing terminals of the A-phase current area. The wiring status is incorrect. It should be connected to the first and third incoming terminals of the A-phase current area; the wiring colors are both yellow, and the wiring colors are correct.

[0233] There is a wire connected to the incoming terminal of the B-phase voltage area, and the wiring status is correct; the wiring color is green, and the wiring color is correct.

[0234] There is a wire connected to the outgoing terminal of the B-phase voltage area, and the wiring status is correct; the wiring color is green, and the wiring color is correct.

[0235] There are two wires connected to the first and third incoming terminals of the B-phase current area. The wiring status is incorrect. It should be connected to the second and third incoming terminals of the B-phase current area; the wiring colors are both green, and the wiring colors are correct.

[0236] There are two wires connected to the first and second outgoing terminals of the B-phase current area. The wiring status is incorrect. It should be connected to the first and third outgoing terminals of the B-phase current area; the wiring colors are both green, and the wiring colors are correct.

[0237] There is a wire connected to the incoming terminal of the C-phase voltage area, and the wiring status is correct; the wiring color is red, and the wiring color is correct.

[0238] There is a wire connected to the outgoing terminal of the C-phase voltage area, and the wiring status is correct; the wiring color is red, and the wiring color is correct.

[0239] There are two wires connected to the first and third incoming terminal positions in the C-phase current area. The wiring status is incorrect. They should be connected to the second and third incoming terminal positions in the B-phase current area. The wiring colors are both red, and the wiring colors are correct.

[0240] There are two wires connected to the first and second outgoing terminal positions in the C-phase current area. The wiring status is incorrect. They should be connected to the first and third outgoing terminal positions in the B-phase current area. The wiring colors are both red, and the wiring colors are correct.

[0241] There is one wire connected to the incoming terminal position in the neutral line area. The wiring status is correct. The wiring color is blue, and the wiring color is correct.

[0242] There is one wire connected to the outgoing terminal position in the neutral line area. The wiring status is correct. The wiring color is blue, and the wiring color is correct.

[0243] }

[0244] 2.2. Marked data review:

[0245] Establish a marked result review mechanism, assign special personnel to conduct strict quality reviews on the marked data to ensure that each mark complies with the specifications. For the discovered marking errors, they should be corrected in a timely manner and feedback to the marking personnel to continuously improve the marking accuracy. At the same time, establish a sample re-review mechanism to conduct sample re-reviews on the marked results to ensure the overall marking quality.

[0246] 3. Model training.

[0247] The present invention selects BLIP-2 as the base model and conducts model training fine-tuning in stages: The training in the first stage is to learn the visual language representation of images based on a frozen image encoder; the training in the second stage is to realize image-to-text generation based on a frozen large language model. After the above two stages of model fine-tuning, the alignment between two different modalities of images and texts is achieved, and the large language model is used to generate the final target text content.

[0248] The BLIP-2 model consists of a pre-trained visual model, the Image Encoder, a pre-trained large language model, the LLM (Large Language Model), and a learnable Q-Former (Query Former). The process of the entire model is as follows: First, the Image Encoder receives an image as input and outputs the visual features of the image; the Q-Former receives the text and the visual features of the image output by the Image Encoder, combines the query vector for fusion analysis, learns visual features similar to the text, and outputs a visual feature representation that the LLM can understand; finally, the LLM receives the visual feature representation output by the Q-Former and generates the corresponding text information. The overall structure of the model is as shown in Figure 6 shown below.

[0249] (1) Image Encoder: A visual model built based on the Vision Transformer architecture, used to extract the visual features of the input image.

[0250] (2) Q-Former: Consists of two sub-modules, the Image Transformer and the Text Transformer, used to bridge the differences between the visual and text modalities and achieve cross-modal alignment.

[0251] (3) LLM: Based on the text generation ability of the large language model, it realizes the ability to generate text from images.

[0252] 3.1 Dataset Preparation:

[0253] Based on the sample format of the standard Image Caption dataset, a training dataset required for fine-tuning the BLIP-2 model is constructed. The dataset includes image data and image description text data. Each image file will have corresponding description text, and these texts will be stored in a JSON (JavaScript Object Notation, a lightweight data interchange format) file. A specific example is as follows:

[0254] ```json

[0256] {

[0257] "image_id (image identifier)": "1",

[0258] "answer": "Junction box description: A junction box appears in the figure. The type of the junction box is a three-phase four-wire one-in-one-out junction box, and the wiring method is a three-phase four-wire wiring method​

[0259] Conclusion on wiring status: Normal wiring status

[0260] Conclusion on wiring color: Normal wiring color

[0261] Detailed description of wiring status and wiring color:

[0262] There is a wire connected to the incoming terminal position in the A-phase voltage area, and the wiring status is correct; the wiring color is yellow, and the wiring color is correct.

[0263] There is a wire connected to the outgoing terminal position in the A-phase voltage area, and the wiring status is correct; the wiring color is yellow, and the wiring color is correct.

[0264] There are two wires connected to the second and third incoming terminal positions in the A-phase current area, and the wiring status is correct; the wiring colors are both yellow, and the wiring colors are correct.

[0265] There are two wires connected to the first and third outgoing terminal positions in the A-phase current area, and the wiring status is correct; the wiring colors are both yellow, and the wiring colors are correct.

[0266] There is a wire connected to the incoming terminal position in the B-phase voltage area, and the wiring status is correct; the wiring color is green, and the wiring color is correct.

[0267] There is a wire connected to the outgoing terminal position in the B-phase voltage area, and the wiring status is correct; the wiring color is green, and the wiring color is correct.

[0268] There are two wires connected to the second and third incoming terminal positions in the B-phase current area, and the wiring status is correct; the wiring colors are both green, and the wiring colors are correct.

[0269] There are two wires connected to the first and third outgoing terminal positions in the B-phase current area, and the wiring status is correct; the wiring colors are both green, and the wiring colors are correct.

[0270] There is a wire connected to the incoming terminal position in the C-phase voltage area, and the wiring status is correct; the wiring color is red, and the wiring color is correct.

[0271] There is a wire connected to the outgoing terminal position in the C-phase voltage area, and the wiring status is correct; the wiring color is red, and the wiring color is correct.

[0272] There are two wires connected to the second and third incoming terminal positions in the C-phase current area, and the wiring status is correct; the wiring colors are both red, and the wiring colors are correct.

[0273] There are two wires connected to the first and third outgoing terminal positions in the C-phase current area, and the wiring status is correct; the wiring colors are both red, and the wiring colors are correct.

[0274] There is a wire connected to the incoming line terminal position in the neutral wire area, and the wiring status is correct; the wiring color is blue, and the wiring color is correct.

[0275] There is a wire connected to the outgoing line terminal position in the neutral wire area, and the wiring status is correct; the wiring color is blue, and the wiring color is correct.

[0276] },

[0277] {

[0278] "image_id":"2",

[0279] "answer":"Description of the junction box: The junction box appears in the figure. The type of the junction box is a three-phase four-wire one-in-one-out junction box, and the wiring method is a three-phase four-wire wiring method.

[0280] Conclusion on the wiring status judgment: Abnormal wiring status

[0281] Conclusion on the wiring color judgment: Normal wiring color

[0282] Detailed description of the wiring status and wiring color:

[0283] There is a wire connected to the incoming line terminal position in the A-phase voltage area, and the wiring status is correct; the wiring color is yellow, and the wiring color is correct.

[0284] There is a wire connected to the outgoing line terminal position in the A-phase voltage area, and the wiring status is correct; the wiring color is yellow, and the wiring color is correct.

[0285] There are two wires connected to the first and third incoming line terminal positions in the A-phase current area. The wiring status is incorrect. They should be connected to the second and third incoming line terminal positions in the A-phase current area; the wiring colors are both yellow, and the wiring color is correct.

[0286] There are two wires connected to the first and second outgoing line terminal positions in the A-phase current area. The wiring status is incorrect. They should be connected to the first and third incoming line terminal positions in the A-phase current area; the wiring colors are both yellow, and the wiring color is correct.

[0287] There is a wire connected to the incoming line terminal position in the B-phase voltage area, and the wiring status is correct; the wiring color is green, and the wiring color is correct.

[0288] There is a wire connected to the outgoing line terminal position in the B-phase voltage area, and the wiring status is correct; the wiring color is green, and the wiring color is correct.

[0289] There are two wires connected to the first and third incoming line terminal positions in the B-phase current area. The wiring status is incorrect. They should be connected to the second and third incoming line terminal positions in the B-phase current area; the wiring colors are both green, and the wiring color is correct.

[0290] There are two wires connected to the first and second outgoing terminal positions in the B-phase current area. The wiring status is incorrect. It should be connected to the first and third outgoing terminal positions in the B-phase current area; the wiring colors are both green, and the wiring colors are correct.

[0291] There is one wire connected to the incoming terminal position in the C-phase voltage area; the wiring status is correct; the wiring color is red, and the wiring color is correct.

[0292] There is one wire connected to the outgoing terminal position in the C-phase voltage area; the wiring status is correct; the wiring color is red, and the wiring color is correct.

[0293] There are two wires connected to the first and third incoming terminal positions in the C-phase current area. The wiring status is incorrect. It should be connected to the second and third incoming terminal positions in the B-phase current area; the wiring colors are both red, and the wiring colors are correct.

[0294] There are two wires connected to the first and second outgoing terminal positions in the C-phase current area. The wiring status is incorrect. It should be connected to the first and third outgoing terminal positions in the B-phase current area; the wiring colors are both red, and the wiring colors are correct.

[0295] There is one wire connected to the incoming terminal position in the neutral line area; the wiring status is correct; the wiring color is blue, and the wiring color is correct.

[0296] There is one wire connected to the outgoing terminal position in the neutral line area; the wiring status is correct; the wiring color is blue, and the wiring color is correct.

[0297] }

[0299] ```

[0300] 3.2. Model Training and Fine-tuning:

[0301] The present invention uses the BLIP-2 model as the base model. This model uses existing pre-trained visual models and large language models, which greatly reduces the model training cost. At the same time, during training, by freezing the parameters of the above two models, catastrophic forgetting of the model is avoided. Therefore, the training and fine-tuning of the BLIP-2 model is actually to fine-tune the Q-Former module, align different modalities of information, and unify them into the feature space that the LLM can understand. It is specifically divided into a two-stage training process:

[0302] (1) The first stage

[0303] ​Based on the frozen parameters of the Image Encoder model, the Q-Former is fully trained using an image-text pair dataset. Meanwhile, three types of losses are designed, namely the contrastive loss (Image-Text Contrastive Learning, ITC), the matching loss (Image-Text Matching, ITM), and the generation loss (Image-Grounded Text Generation, ITG), to optimize the parameters of the Q-Former model. The specific calculation of the loss function is as follows:

[0304] The matching loss is used to evaluate whether an image and text belong to the same instance, that is, to determine whether the image and text match, which belongs to a binary classification task. For each pair of image-text (I i , T i ), based on the Q-Former, a matching score p match (I i , T i ) is generated, which is the probability that the image and text match. Its matching loss L ML is calculated as follows:

[0305]

[0306] where N is the number of samples; y i is the binary label, taking the value of 1 indicating that (I i , T i ) is a matching pair, and taking the value of 0 indicating a non-matching pair.

[0307] The generation loss is used to evaluate the model's ability to generate text, usually involving the task of a masked language model. Part of the text is masked, and the model needs to predict the masked text. Given a text sequence T and a masked sequence T′, for the position t in the text, if the position is masked, the probability of predicting the word ω t at this position can be calculated as p(ω t |T′). Then the generation loss L GL is specifically calculated as follows:

[0308] L GL = -∑ t∈masked positions logp(ω t |T′), Equation (2)

[0309] The contrastive loss is to make the distance between the embedding vectors of a pair of matching images and texts as small as possible, and the distance between different pairs as large as possible. The contrastive loss L CL is specifically calculated as follows:

[0310]

[0311] where N is the number of samples; s(I i , T i ) represents the similarity score of the image-text pair (I i , T i ); τ is the temperature parameter used to adjust the similarity score distribution, and generally takes values between 0.07 and 0.1.

[0312] The ultimate goal is to transform the original image features extracted by the Image Encoder into visual features related to text features through learnable Queries (queries). The model training structure is as Figure 7 shown.

[0313] (2) The second stage

[0314] Based on the frozen LLM model parameters, the image features transformed by the Q-Former are input into the fully connected layer, mapped to the text embedding space of the LLM, combined with image information and text information, and finally the target text content is generated by the LLM. The specific model structure is as Figure 8 shown.

[0315] 4. Model deployment and service encapsulation.

[0316] 4.1 Model deployment:

[0317] According to the requirements of the actual application and the computing power resources required for model inference, select a target system with suitable GPU (Graphics Processing Unit) inference computing power resources, install the environmental software packages required for model inference, ensure that the model can perform model inference normally on the target system, load the fine-tuned model parameters, and perform model deployment.

[0318] 4.2 Model service encapsulation:

[0319] To enhance the application flexibility and convenience of the model, the present invention encapsulates the use of the model into an API (Application Programming Interface) interface, enabling it to be widely applied to different systems and platforms. The API interface follows the RESTful (Representational State Transfer) principle and returns the wiring status of the junction box in the image by receiving the image of the junction box and the corresponding text input.

[0320] It should be noted that in this embodiment, BLIP-2 is used as the basic multimodal large model, which can be replaced by other multimodal large models and fine-tuning frameworks.

[0321] In the above embodiment, when identifying the wiring state of the energy metering junction box, the multimodal large model is pre-trained by using the sample image data and sample text data of the sample energy metering junction box, so that in practical applications, after obtaining the image data of the energy metering junction box to be analyzed, the wiring state recognition result of the energy metering junction box to be analyzed can be predicted; secondly, the multimodal large model receives new data in each round of iterative training, so as to improve and optimize the internal model, so as to be able to predict more effectively, which is conducive to improving the prediction accuracy of the multimodal large model, and thus improving the recognition accuracy of the wiring state of the energy metering junction box; moreover, the whole process does not require manual intervention, avoiding the subjective factors in the method of artificial visual detection, which is prone to errors and leads to the defect of low recognition accuracy of the wiring state of the energy metering junction box, and further improves the recognition accuracy of the wiring state of the energy metering junction box. At the same time, in view of the problem that the existing AI recognition system has high requirements for the image quality of the junction box, the present invention effectively overcomes the influence of factors such as ambient light and shooting angle on the recognition accuracy by introducing a multimodal large model. The multimodal large model combines the capabilities of image recognition and text processing, and can make comprehensive use of image data and related text information (such as operating manuals, maintenance records, etc.) to improve the robustness of the model to image quality fluctuations. Even in the case of blur, poor lighting, complex background or improper shooting angle, the model can still accurately identify the wiring status, significantly improving the accuracy of recognition. For the existing AI model requires a large amount of labeled data for training, and the high cost of obtaining high-quality labeled data, the present invention reduces the dependence on a large amount of high-quality labeled data by adopting a multimodal learning strategy. The BLIP-2 model can effectively utilize unlabeled data or low-quality labeled data, and improves learning efficiency and model performance through the complementarity of image and text information during model training. This strategy greatly reduces the cost of data annotation, making the popularization and application of AI recognition systems more feasible. In view of the problem that the current mainstream image recognition algorithm can only determine whether there is a wiring abnormality, but cannot clearly point out the abnormal location and give correct wiring suggestions, the present invention realizes the precise positioning and suggestion prompts of wiring abnormalities through the comprehensive application of the BLIP-2 model. It can not only determine whether the wiring is abnormal, but also point out the specific abnormal location through the fusion analysis of image and text information, and give suggestions for correct wiring in combination with text information such as the operation manual. This progress has greatly improved the efficiency and accuracy of power system maintenance and safety management.

[0322] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0323] Based on the same inventive concept, an embodiment of the present application further provides a multimodal large model-based power metering junction box wiring state recognition device for implementing the above-mentioned multimodal large model-based power metering junction box wiring state recognition method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the multimodal large model-based power metering junction box wiring state recognition device provided below can refer to the limitations on the multimodal large model-based power metering junction box wiring state recognition method in the above text, and will not be repeated here.

[0324] In an exemplary embodiment, as Figure 9 shown, a multimodal large model-based power metering junction box wiring state recognition device is provided, including: a data acquisition module 901, a data determination module 902, a loss determination module 903, a target determination module 904, a model training module 905, and a result determination module 906, where:

[0325] The data acquisition module 901 is configured to acquire sample image data of a sample power metering junction box.

[0326] The data determination module 902 is configured to determine, through a multimodal large model to be trained, the image visual features corresponding to the sample image data, determine the text visual features corresponding to the sample image data according to the image visual features, and determine the sample text data corresponding to the sample image data according to the text visual features.

[0327] The loss determination module 903 is configured to determine a matching loss value, a generation loss value, and a contrast loss value according to the multimodal large model to be trained, the sample image data, and the sample text data.

[0328] The target determination module 904 is configured to obtain a target loss value according to the matching loss value, the generation loss value, and the contrast loss value.

[0329] The model training module 905 is used to iteratively train the multi-modal large model to be trained according to the target loss value, and obtain the trained multi-modal large model.

[0330] The result determination module 906 is used to obtain the image data of the electricity metering junction box to be analyzed, and input the image data into the trained multi-modal large model to obtain the wiring status recognition result of the electricity metering junction box to be analyzed.

[0331] In an exemplary embodiment, the data determination module 902 is further used to determine the image visual features corresponding to the sample image data through an image encoder; determine the text visual features corresponding to the sample image data according to the image visual features through a query transformer; and determine the sample text data corresponding to the sample image data according to the text visual features through a large language model.

[0332] In an exemplary embodiment, the loss determination module 903 is further used to input the sample image data and the sample text data into the multi-modal large model to be trained, and obtain the matching probability between the sample image data and the sample text data; and determine the matching loss value according to the matching probability.

[0333] In an exemplary embodiment, the loss determination module 903 is further used to perform masking processing on the sample text data to obtain the processed sample text data carrying the masked text data; input the sample image data and the processed sample text data into the multi-modal large model to be trained, and obtain the prediction probability corresponding to the masked text data in the processed sample text data; and determine the generation loss value according to the prediction probability.

[0334] In an exemplary embodiment, the loss determination module 903 is further used to determine the first embedding vector corresponding to the sample image data and the second embedding vector corresponding to the sample text data through the multi-modal large model to be trained; determine the similarity between the first embedding vector and the second embedding vector; and determine the contrast loss value according to the similarity.

[0335] In an exemplary embodiment, the target determination module 904 is further used to obtain the first preset weight corresponding to the matching loss value, the second preset weight corresponding to the generation loss value, and the third preset weight corresponding to the contrast loss value; and perform weighted summation processing on the matching loss value, the generation loss value, and the contrast loss value according to the first preset weight, the second preset weight, and the third preset weight to obtain the target loss value.

[0336] In an exemplary embodiment, the data acquisition module 901 is further configured to acquire system image data, on-site image data, and simulated image data associated with the sample power metering junction box; preprocess the system image data, on-site image data, and simulated image data to obtain preprocessed system image data, preprocessed on-site image data, and preprocessed simulated image data; and use the preprocessed system image data, preprocessed on-site image data, and preprocessed simulated image data as the sample image data of the sample power metering junction box.

[0337] Each module in the above power metering junction box wiring state recognition device based on the multi-modal large model can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0338] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample image data, sample text data, etc. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for recognizing the wiring state of a power metering junction box based on a multi-modal large model.

[0339] Those skilled in the art can understand that Figure 10 the structure shown in

[0340] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0341] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0342] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0343] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0344] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0345] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for identifying the wiring status of an electric energy metering junction box based on a multimodal large model, characterized in that: The method comprises: Acquire sample image data of a sample electric energy metering junction box; Determine the image visual features corresponding to the sample image data through the multimodal large model to be trained, determine the text visual features corresponding to the sample image data based on the image visual features, and determine the sample text data corresponding to the sample image data based on the text visual features; Determine a matching loss value, a generation loss value, and a comparison loss value according to the multimodal large model to be trained, the sample image data, and the sample text data; Obtaining a target loss value according to the matching loss value, the generation loss value and the comparison loss value; Iteratively training the multimodal large model to be trained according to the target loss value to obtain a trained multimodal large model; Image data of the electric energy metering junction box to be analyzed is obtained, and the image data is input into the trained multimodal large model to obtain a wiring state recognition result of the electric energy metering junction box to be analyzed.

2. The method according to claim 1, characterized in that The multimodal large model to be trained includes an image encoder, a query transformer and a large language model; The method of determining the image visual features corresponding to the sample image data by using the multimodal large model to be trained, determining the text visual features corresponding to the sample image data according to the image visual features, and determining the sample text data corresponding to the sample image data according to the text visual features includes: Determining, by means of the image encoder, image visual features corresponding to the sample image data; Determining, by the query transformer, text visual features corresponding to the sample image data according to the image visual features; The sample text data corresponding to the sample image data is determined by the large language model according to the text visual features.

3. The method according to claim 1, characterized in that The determining of the matching loss value, the generation loss value and the comparison loss value according to the multimodal large model to be trained, the sample image data and the sample text data includes: Inputting the sample image data and the sample text data into the multimodal large model to be trained to obtain a matching probability between the sample image data and the sample text data; The matching loss value is determined according to the matching probability.

4. The method according to claim 1, characterized in that: The method of determining a matching loss value, a generation loss value and a comparison loss value according to the multimodal large model to be trained, the sample image data and the sample text data further includes: Performing mask processing on the sample text data to obtain processed sample text data carrying the masked text data; Inputting the sample image data and the processed sample text data into the multimodal large model to be trained to obtain the prediction probability corresponding to the mask text data in the processed sample text data; The generation loss value is determined based on the predicted probability.

5. The method according to claim 1, characterized in that The method of determining a matching loss value, a generation loss value and a comparison loss value according to the multimodal large model to be trained, the sample image data and the sample text data further includes: Determine, by means of the multimodal large model to be trained, a first embedding vector corresponding to the sample image data and a second embedding vector corresponding to the sample text data; Determining a similarity between the first embedding vector and the second embedding vector; The contrast loss value is determined according to the similarity.

6. The method according to claim 1, characterized in that The obtaining a target loss value according to the matching loss value, the generated loss value and the contrast loss value includes: Obtaining a first preset weight corresponding to the matching loss value, a second preset weight corresponding to the generation loss value, and a third preset weight corresponding to the comparison loss value; According to the first preset weight, the second preset weight and the third preset weight, the matching loss value, the generation loss value and the comparison loss value are weighted and summed to obtain the target loss value.

7. The method according to any one of claims 1 to 6, characterized in that The step of obtaining sample image data of a sample electric energy metering junction box includes: Acquire system image data, field image data, and simulation image data associated with a sample electric energy metering junction box; Preprocessing the system image data, the field image data and the simulation image data to obtain preprocessed system image data, preprocessed field image data and preprocessed simulation image data; The preprocessed system image data, the preprocessed field image data and the preprocessed simulation image data are all used as sample image data of the sample electric energy metering junction box.

8. A device for identifying the wiring status of an electric energy metering junction box based on a multimodal large model, characterized in that: The device comprises: A data acquisition module, used to acquire sample image data of a sample electric energy metering junction box; A data determination module, used to determine the image visual features corresponding to the sample image data through the multimodal large model to be trained, determine the text visual features corresponding to the sample image data according to the image visual features, and determine the sample text data corresponding to the sample image data according to the text visual features; A loss determination module, used to determine a matching loss value, a generation loss value and a comparison loss value according to the multimodal large model to be trained, the sample image data and the sample text data; a target determination module, configured to obtain a target loss value according to the matching loss value, the generation loss value and the comparison loss value; A model training module, used for iteratively training the multimodal large model to be trained according to the target loss value to obtain a trained multimodal large model; The result determination module is used to obtain image data of the electric energy metering junction box to be analyzed, and input the image data into the trained multi-modal large model to obtain the wiring state recognition result of the electric energy metering junction box to be analyzed.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.