Grid Dispatching Automation Intelligent Operation and Maintenance Method and System Based on Multimodal Large Model

By constructing a multimodal large model to process text and image information, and combining the search and verification of the historical operation and maintenance solution to confirm the reliability of the output results, the problem that the AIGC model operation and maintenance system cannot judge the output results by itself, and improves the reliability and efficiency of operation and maintenance.

CN119379263BActive Publication Date: 2025-06-27STATE GRID ZHEJIANG ELECTRIC POWER CO LTD +1

Patent Information

Application Number
CN202411921274.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-06-27
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

In the prior art, intelligent operation and maintenance systems based on the AIGC model cannot determine whether the output results are reliable by themselves, resulting in difficulty in ensuring operation and maintenance effects and efficiency.

Method used

By building a multimodal large model, including a language processing model, image recognition model and AIGC operation and maintenance model, input text and image information for processing, generate operation and maintenance plan text, and conduct three-party confirmation through the keyword search of the historical operation and maintenance plan to judge the matching degree to determine the reliability of the output results.

Benefits of technology

It realizes the reliability of the operation and maintenance system being able to judge and prompt the results when outputting results, reduces the review work of operation and maintenance personnel, and greatly improves the reliability and efficiency of operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119379263B_ABST
    Figure CN119379263B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for intelligent operation and maintenance of power grid dispatching automation based on a multi-modal large model. A multi-modal large model composed of a language processing model, an image recognition model, and an AIGC operation and maintenance model is constructed. By processing the input information from different information sources, different operation and maintenance plan texts are obtained. Then, historical operation and maintenance plans are retrieved based on the keywords of the input information for tripartite confirmation. When the matching degree is lower than the threshold, the output result is considered unreliable. Through innovations such as multi-modal information input, dual operation and maintenance plan output, and matching degree judgment, the present invention solves the problem that generative AI is prone to misrepresentation, enables the operation and maintenance system to judge and prompt the reliability of the output result when outputting results, eliminates the frequent review by operation and maintenance personnel, greatly improves the reliability and efficiency of operation and maintenance, realizes a comprehensive understanding and efficient processing of the operation and maintenance scenario, and provides a strong technical guarantee for the safe and stable operation of the power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a power grid dispatching automation intelligent operation and maintenance method and system based on a multi-modal large model. Background Art

[0002] With the rapid development of technologies such as 5G, cloud computing, and the Internet of Things, the scale and complexity of networks have increased rapidly. The traditional manual operation and maintenance mode is gradually difficult to cope with a large number of devices and complex network problems, and has low operation and maintenance efficiency, lagging fault handling, and continuously rising operation and maintenance costs. Based on this demand, AI (Artificial Intelligence) technology has begun to be gradually introduced into the operation and maintenance field, forming the concept of intelligent operation and maintenance (AIOps). Existing technologies have proposed various architectures. For example, an operation and maintenance intelligent solution that combines technologies such as AI, big data, and RPA (Robotic Process Automation), through the combination of multi-modal input, intent understanding, knowledge graph, and AIGC (Generative AI) model, creates "operation and maintenance digital employees" to improve the intelligent level of network operation and maintenance.

[0003] Thanks to the emergence of the AIGC model, the output results of various intelligent operation and maintenance solutions are becoming more and more accurate, providing a new path for improving operation and maintenance efficiency. However, due to the large number of knowledge points involved in operation and maintenance, and problems such as inaccurate information transmission that may occur during the feature transformation process after multi-modal input, the intelligent operation and maintenance solution based on the AIGC model will still have occasional output results that do not conform to reality or other errors. And the most serious problem is that the model usually only outputs content step by step, and it cannot judge whether the result is reliable by itself. Therefore, for operation and maintenance personnel, if they want to ensure the reliability of operation and maintenance, they must manually review each output result, which reduces the efficiency gap between intelligent operation and maintenance and manual operation and maintenance, and also makes intelligent operation and maintenance meaningless.

[0004] Therefore, how to enable the operation and maintenance system to be aware of its own limitations and prompt risks in a timely manner when it cannot accurately judge is the key to improving operation and maintenance efficiency and also a technical problem that is difficult to solve at present. Summary of the Invention

[0005] In view of the problem that the existing operation and maintenance system cannot judge whether the output result is reliable by itself, resulting in difficulties in ensuring the operation and maintenance effect and efficiency, the present invention provides a power grid dispatching automation intelligent operation and maintenance method and system based on a multi-modal large model. By processing the input information from different information sources, different operation and maintenance plan texts are obtained, and then historical operation and maintenance plans are retrieved based on the keywords of the input information for tripartite confirmation. When the matching degree is lower than the threshold, the output result is considered unreliable. The present invention solves the problem that generative AI is prone to misrepresentation through tripartite confirmation, enables the operation and maintenance system to judge and prompt the reliability of the result when outputting the result, eliminates the frequent review by operation and maintenance personnel, and greatly improves the operation and maintenance reliability and efficiency.

[0006] The following is the technical solution of the present invention.

[0007] A power grid dispatching automation intelligent operation and maintenance method based on a multi-modal large model includes the following steps:

[0008] S1: Pre-construct a multi-modal large model composed of a language processing model, an image recognition model, and an AIGC operation and maintenance model;

[0009] S2: Input text input information and image input information, and superimpose marking information on the image input information;

[0010] S3: The image recognition model in the multi-modal large model recognizes the image input information and the marking information, and outputs text description information;

[0011] S4: The language processing model in the multi-modal large model compares and synthesizes the text input information and the text description information to generate a collection of text clues;

[0012] S5: The AIGC operation and maintenance model in the multi-modal large model independently reads the text input information and the collection of text clues respectively, outputs corresponding different operation and maintenance plan texts, and at the same time retrieves historical operation and maintenance plans based on the keywords of the collection of text clues to obtain retrieval results;

[0013] S6: Judge the matching degree between different operation and maintenance plan texts and the retrieval results. If the matching degrees are all lower than the threshold, execute S7; otherwise, output the operation and maintenance plan texts and retrieval results with a matching degree not lower than the threshold;

[0014] S7: Output the operation and maintenance plan text and the retrieval result, and prompt risk information. If a change instruction is received, obtain the changed text input information and / or text description information based on the change instruction, and return to S4. If a reset instruction is received, return to S2.

[0015] In the present invention, a multi-modal large model composed of different sub-models is used to support the input and processing of text and images. Among them, the text input is directly converted into text input information, while the image input information and the marking information are converted into text description information. The text input information and the text description information are used to generate a collection of text clues to express the complete input content. Subsequently, the present invention independently reads different operation and maintenance plan texts obtained from the text input information and the collection of text clues respectively, and at the same time retrieves the historical operation and maintenance plans according to the collection of text clues. Since one of the operation and maintenance plan texts only considers the text input information, while the other operation and maintenance plan text considers both the text input information and the text description information (indirectly considering the image content), the reliability of the two operation and maintenance plan texts is in doubt when it is uncertain whether the expression of the text or the image is clear. Therefore, the present invention continues to verify by comparing the historical operation and maintenance plans, and can determine the reliability of the operation and maintenance plan texts, and clearly prompt the risks while displaying. On the one hand, it can be used as a reference for operation and maintenance personnel, and on the other hand, it does not require operation and maintenance personnel to verify the reliability of each result, improving the operation and maintenance efficiency.

[0016] Preferably, in step S1: a multi-modal large model composed of a language processing model, an image recognition model, and an AIGC operation and maintenance model is pre-constructed, including:

[0017] Based on a number of Transformer encoders, a BERT model is constructed, and the BERT model is trained to obtain a language processing model;

[0018] Based on a real-time object detection algorithm, a YOLO model is constructed, and the YOLO model is trained to obtain an image recognition model;

[0019] Based on the architecture of Transformer, a GPT model is constructed by using a multi-layer self-attention mechanism and position encoding, and the GPT model is trained to obtain an AIGC operation and maintenance model.

[0020] Preferably, in step S2: the text input information and the image input information are input, and the marking information is superimposed on the image input information, including:

[0021] The text input information is input by means of voice or text input;

[0022] The image input information is input by means of shooting or importing;

[0023] The image input information is displayed, and the operator draws a closed figure in the image input information to form the marking information, records the coordinates of the marking information, and superimposes the marking information on the image input information based on the coordinate position.

[0024] Preferably, in step S3: the image recognition model in the multi-modal large model recognizes the image input information and the marking information, and outputs the text description information, including:

[0025] The image recognition model in the multi-modal large model performs object detection on the image input information, identifies the scene in the image, outputs scene description information, and performs separate object detection on the image area involved in the marking information to identify and output device description information about the target device in the image area;

[0026] Based on a preset description template, integrate the scene description information and the device description information to output text description information.

[0027] Preferably, in step S4: The language processing model in the multi-modal large model compares and synthesizes the text input information and the text description information to generate a collection of text clues, including:

[0028] The language processing model in the multi-modal large model identifies the text input information and the text description information to obtain keywords related to the operation and maintenance tasks as clue information, and represents the extracted clue information in a structured manner;

[0029] Perform alignment and fusion processing on the clue information to obtain a collection of text clues.

[0030] Preferably, in step S5: The AIGC operation and maintenance model in the multi-modal large model independently reads the text input information and the collection of text clues respectively, and outputs corresponding different operation and maintenance plan texts. At the same time, retrieve the historical operation and maintenance plans based on the keywords in the collection of text clues to obtain retrieval results, including:

[0031] The AIGC operation and maintenance model in the multi-modal large model reads the text input information and outputs the first operation and maintenance plan text based on the text input information;

[0032] The AIGC operation and maintenance model in the multi-modal large model reads the collection of text clues and outputs the second operation and maintenance plan text based on the collection of text clues;

[0033] At the same time, retrieve the historical operation and maintenance plans, and search for the keywords in the collection of text clues in the historical operation and maintenance plans. The operation and maintenance plan corresponding to the training set with the highest relevance is used as the retrieval result.

[0034] In the present invention, the text input information and the text description information respectively correspond to the information contents carried by the text input and the picture input in the operation and maintenance scenario, while the text clue collection carries all the input contents in the current operation and maintenance scenario. Since it cannot be ensured whether the content is complete, nor can it be ensured whether each model understands the meaning therein, two different operation and maintenance plan texts are respectively output, and the historical operation and maintenance plans are retrieved by keywords as a reference. Only when the description of the operation and maintenance plan text has a high degree of matching with the historical operation and maintenance plan is the output result considered credible, solving the problem that the existing generative AI model cannot evaluate whether the output content is credible, and at the same time avoiding the problems of information loss and incomplete scenario understanding that may be caused by only retrieving by keywords.

[0035] Preferably, step S6: judging the matching degrees of different operation and maintenance plan texts and the retrieval results. If the matching degrees are all lower than the threshold, then execute S7; otherwise, output the operation and maintenance plan text and the retrieval results with the matching degree not lower than the threshold, including:

[0036] Determine the judgment dimensions of the matching degrees, and compare different operation and maintenance plan texts with the retrieval results based on the judgment dimensions to obtain the matching degrees of each operation and maintenance plan text and the retrieval results;

[0037] If the matching degrees are all lower than the threshold, then execute S7;

[0038] Otherwise, output the operation and maintenance plan text and the retrieval results with the matching degree not lower than the threshold, and this operation and maintenance plan text and the retrieval results will be recognized as reliable plans.

[0039] Preferably, step S7: output the operation and maintenance plan text and the retrieval results, and prompt risk information. If a change instruction is received, then obtain the changed text input information and / or text description information based on the change instruction, and return to S4. If a reset instruction is received, then return to S2, including:

[0040] Use different colors to mark the same parts and different parts in the operation and maintenance plan text and the retrieval results, and use the prompt of the risk information to cover the display areas of the operation and maintenance plan text and the retrieval results;

[0041] After the confirmation button in the risk information is clicked, clear the risk information and redisplay the operation and maintenance plan text and the retrieval results;

[0042] If it is selected to change the text input information and / or text description information, then display the text input information and / or text description information completely and set it to be editable. After editing based on the change instruction, generate the changed text input information and / or text description information;

[0043] If a reset instruction is received, then return to S2.

[0044] The present invention also provides a power grid dispatching automation intelligent operation and maintenance system based on a multi-modal large model, including an operation and maintenance server and an operation and maintenance terminal, and the operation and maintenance server and the operation and maintenance terminal are configured to execute the above-mentioned power grid dispatching automation intelligent operation and maintenance method based on the multi-modal large model.

[0045] The present invention also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the above-mentioned power grid dispatching automation intelligent operation and maintenance method based on the multi-modal large model are realized.

[0046] The present invention also provides a storage medium, wherein computer-executable instructions are stored in the storage medium, and when the computer-executable instructions are loaded and executed by a processor, the steps of the above-mentioned power grid dispatching automation intelligent operation and maintenance method based on the multi-modal large model are realized.

[0047] The substantial effects of the present invention include:

[0048] By constructing a multi-modal large model composed of a language processing model, an image recognition model, and an AIGC operation and maintenance model, the present invention realizes the comprehensive processing ability of text and image information. This multi-modal input method enables operation and maintenance personnel to input information in a diversified manner, improving the flexibility and accuracy of information input. At the same time, through the image recognition model, target detection is performed on the image input information, and the key information in the image is converted into text description information, further enriching the description of the operation and maintenance scenario and providing a more comprehensive information basis for the subsequent generation of operation and maintenance plans.

[0049] The present invention outputs two different operation and maintenance plan texts by separately and independently reading the text input information and the text clue collection synthesized from the text input information and the text description information. At the same time, in combination with the retrieval results of historical operation and maintenance plans based on the keywords in the text clue collection, through a matching degree judgment mechanism, the consistency between the operation and maintenance plan text and the retrieval results is evaluated, which can further verify the reliability and feasibility of the operation and maintenance plan. When the matching degree is lower than the threshold, the system will prompt risk information and allow operation and maintenance personnel to perform modification or reset operations as needed. This solves the problem that generative AI in the prior art may "talk nonsense", thereby improving the accuracy of operation and maintenance decisions.

[0050] Finally, the present invention enhances the readability and comprehensibility of the operation and maintenance plan by using different colors to mark the same and different parts in the operation and maintenance plan text and the retrieval results, and by covering the display area with the prompt of risk information. These user-friendly designs enable operation and maintenance personnel to grasp key information more quickly, improving the operation and maintenance efficiency.

[0051] In summary, through innovative features such as multi-modal information input, dual operation and maintenance solution output, matching degree judgment, and user-friendly display methods, the present invention realizes a comprehensive understanding and efficient processing of the operation and maintenance scenarios, improves the operation and maintenance efficiency and accuracy, and provides a strong technical guarantee for the safe and stable operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in combination with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0054] It should be understood that in various embodiments of the present invention, the sequence numbers of the processes do not imply the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0055] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0056] It should be understood that in the present invention, "a plurality of" means two or more. "And / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. "Including A, B, and C" and "including A, B, C" mean that all of A, B, and C are included. "Including A, B, or C" means including one of A, B, and C. "Including A, B, and / or C" means including any one or any two or all three of A, B, and C.

[0057] The following will detail the technical solutions of the present invention with specific embodiments. The embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0058] Embodiment 1: A method for intelligent operation and maintenance of power grid dispatching automation based on a multi-modal large model, such asFigure 1 As shown, it includes the following steps:

[0059] S1: Pre-construct a multi-modal large model composed of a language processing model, an image recognition model, and an AIGC operation and maintenance model.

[0060] It includes: constructing a BERT model based on several Transformer encoders, and training the BERT model to obtain a language processing model;

[0061] Constructing a YOLO model based on a real-time object detection algorithm, and training the YOLO model to obtain an image recognition model;

[0062] Based on the Transformer architecture, constructing a GPT model using a multi-layer self-attention mechanism and positional encoding, and training the GPT model to obtain an AIGC operation and maintenance model.

[0063] In this embodiment, language processing is performed by constructing a BERT model. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language representation model based on Transformer encoders, and its core lies in capturing context information in text through a multi-layer bidirectional Transformer structure. During the construction process, first design and initialize the parameters of the Transformer encoder, including the dimensions of the input layer, hidden layer, and output layer. Then use a large amount of text data to train the BERT model so that it can deeply understand the context and meaning of the language. After training, this model can be used for various language processing tasks, such as various semantic analyses related to operation and maintenance.

[0064] Secondly, for image recognition, this embodiment uses the YOLO (You Only Look Once) model. YOLO is a real-time object detection algorithm that can quickly and accurately identify various objects in a picture. The network structure of YOLO includes convolutional layers, pooling layers, and fully connected layers, etc., to extract image features and predict the target position and category. Then, use the labeled image dataset for training, and optimize the model parameters through the backpropagation algorithm and loss function, so that the model can accurately identify the targets in the image. By training the YOLO model, it can be made to have powerful image recognition capabilities, so as to identify scenes, target devices, etc. during the operation and maintenance process.

[0065] Finally, to support AIGC operation and maintenance, this embodiment constructs a GPT (Generative Pre-trained Transformer) model. GPT is a generative pre-trained model based on the Transformer architecture, which uses multi-layer self-attention mechanisms and positional encoding to generate coherent text content. During the construction process, first, the parameters of the Transformer decoder are designed and initialized, including the dimensions of the input layer, hidden layer, and output layer. Then, a large-scale text corpus is used for pre-training. The training samples in this embodiment are mainly historical operation and maintenance solutions, etc. The model is trained through an autoregressive language modeling task to enable it to learn to generate text that conforms to language rules. During the training process, an adaptive learning rate and optimization algorithm are used to update the model parameters until the model achieves the best performance on the validation set. By training the GPT model, it can be enabled to have the ability to generate high-quality text content, thereby supporting AIGC operation and maintenance tasks.

[0066] S2: Input the text input information and the image input information, and superimpose the marking information on the image input information.

[0067] Including: inputting the text input information by means of voice or text input;

[0068] inputting the image input information by means of shooting or importing;

[0069] Display the image input information, and the operator draws a closed figure in the image input information to form the marking information, records the coordinates of the marking information, and superimposes the marking information on the image input information based on the coordinate position.

[0070] In this embodiment, the system receives text input, which can be achieved through speech-to-text technology or direct text input. For example, voice inputs such as "The oil level gauge reading of XX transformer is low", "The running sound of XX equipment is abnormal and the vibration is too large", etc.

[0071] At the same time, the image input is completed by taking new photos or importing pictures from existing resources, which ensures the wide range of data sources. In the image display interface of this embodiment, the operator can intuitively see the image content and accurately mark the key areas or objects in the image by drawing closed figures (such as rectangles, circles, etc.). These markings not only help to clarify the focus of attention but also provide accurate positioning information for subsequent analysis or processing. Generally speaking, the marked and circled parts are mainly the target devices, and this method can facilitate the system to grasp the key content in subsequent processing.

[0072] To achieve this, the system records the coordinate data of each marker, which is an exact position indication of the marker information on the image. Based on these coordinates, the system can accurately overlay the marker information on the original image, forming a rich image dataset with an additional annotation layer.

[0073] S3: The image recognition model in the multimodal large model recognizes the image input information and the marker information, and outputs text description information.

[0074] Including: The image recognition model in the multimodal large model performs object detection on the image input information, recognizes the scene in the image, outputs scene description information, and performs separate object detection on the image area involved in the marker information, recognizes and outputs device description information about the target device in the image area;

[0075] Based on a preset description template, the scene description information and the device description information are integrated to output text description information.

[0076] In this embodiment, the image recognition model in the multimodal large model performs comprehensive object detection on the image input information. This process involves the recognition of various elements in the image, including people, objects, backgrounds, etc., so as to generate description information about the overall scene. Such a description is not limited to the prominent or main elements in the image, but also covers the overall environment and atmosphere presented by the image. For example, the image recognition model in the multimodal large model performs comprehensive object detection on the picture, recognizes the main elements in the picture, including transformers, fences, utility poles, etc., and generates description information about the overall scene, such as "The picture shows a power grid operation and maintenance scene, with a large transformer in the center, surrounded by a fence, and there are several utility poles nearby".

[0077] Then, the model performs more refined object detection on the image area specified by the marker information. This usually involves the recognition of specific devices, items, or specific parts in the image, and outputs detailed description information about these target devices. This targeted detection and analysis make the output description information more accurate and specific. For example, the marker information marks the area of the transformer, and the model performs more refined object detection on this specific device, and outputs detailed description information about the transformer, such as "The transformer model is XX, the current oil level is abnormal, there may be signs of oil leakage, and the radiator is in good condition".

[0078] Finally, based on a preset description template, the model integrates the scene description information and the device description information. This step ensures that the output text description information is both comprehensive and structured, facilitating quick understanding and use by users. The design of the description template can be adjusted according to the actual application scenario to meet the requirements of different fields and uses. For example, based on the preset description template, the text description information is output as follows: "Overall environment: A transformer is surrounded by a fence, and there are several utility poles nearby. The overall environment is safe; Target device: A transformer of model XX; Abnormality judgment: Abnormal oil level, no abnormality in the radiator, current status is normal."

[0079] S4: The language processing model in the multimodal large model compares and synthesizes the text input information and the text description information to generate a collection of text clues.

[0080] It includes: The language processing model in the multimodal large model identifies the text input information and the text description information to obtain keywords related to the operation and maintenance tasks as clue information, and represents the extracted clue information in a structured manner;

[0081] Perform alignment and fusion processing on the clue information to obtain a collection of text clues.

[0082] For example, if the text input information mentions "equipment overheating" and the text description information identifies "equipment damage" from the picture, then after fusion, it can be obtained as "equipment damage, overheating". Alignment is considered to ensure that when the information is complex, the description objects of the equipment need to be aligned to avoid statements describing different equipment from being integrated into the description of one equipment.

[0083] S5: The AIGC operation and maintenance model in the multimodal large model independently reads the text input information and the collection of text clues respectively, outputs corresponding different operation and maintenance plan texts, and at the same time retrieves the historical operation and maintenance plans based on the keywords in the collection of text clues to obtain the retrieval results.

[0084] It includes: The AIGC operation and maintenance model in the multimodal large model reads the text input information and outputs the first operation and maintenance plan text based on the text input information;

[0085] The AIGC operation and maintenance model in the multimodal large model reads the collection of text clues and outputs the second operation and maintenance plan text based on the collection of text clues;

[0086] At the same time, retrieve the historical operation and maintenance plans, and search for the keywords in the collection of text clues in the historical operation and maintenance plans. The operation and maintenance plan corresponding to the training set with the highest relevance is used as the retrieval result.

[0087] In this embodiment, the text input information and the text description information respectively correspond to the information contents carried by the text input and the picture input in the operation and maintenance scenario, while the text clue collection carries all the input contents in the current operation and maintenance scenario. Since it cannot be ensured whether the content is complete or whether each model understands the meaning therein, two different operation and maintenance solution texts are output respectively, and the historical operation and maintenance solutions are retrieved by keywords as a reference. Only when the description of the operation and maintenance solution text has a high degree of match with the historical operation and maintenance solution is the output result considered credible, solving the problem that the existing generative AI model cannot evaluate whether the output content is credible, and at the same time avoiding the problems of information loss and incomplete scenario understanding that may be caused by keyword retrieval alone.

[0088] S6: Judge the matching degrees of different operation and maintenance solution texts and the retrieval results. If the matching degrees are all lower than the threshold, execute S7; otherwise, output the operation and maintenance solution text and the retrieval results whose matching degrees are not lower than the threshold.

[0089] It includes: determining the judgment dimension of the matching degree, comparing different operation and maintenance solution texts with the retrieval results based on the judgment dimension, and obtaining the matching degree of each operation and maintenance solution text and the retrieval results;

[0090] If the matching degrees are all lower than the threshold, execute S7;

[0091] Otherwise, output the operation and maintenance solution text and the retrieval results whose matching degrees are not lower than the threshold, and this operation and maintenance solution text and the retrieval results will be recognized as reliable solutions.

[0092] In this embodiment, the semantic similarity is selected as the judgment dimension of the matching degree. When judging, the text or vocabulary is represented as a vector in a high-dimensional space, and then the similarity between these vectors (such as cosine similarity, Euclidean distance, etc.) is calculated. By calculating the similarity between two vocabulary vectors, the semantic similarity between them can be evaluated, and then the matching degree can be obtained.

[0093] This embodiment can avoid various situations where the output results are unreliable. For example, keywords may not accurately reflect the actual situation, and the retrieved historical operation and maintenance solutions are not suitable. Since the two different operation and maintenance solution texts consider the subjective expressions of the operation and maintenance personnel, the obtained content does not match the retrieval results. At this time, this embodiment will prompt risks. Another example is that there are misjudgments or ambiguities in the subjective expressions of the operation and maintenance personnel, resulting in misunderstandings in the two different operation and maintenance solution texts, while the historical operation and maintenance solutions retrieved by keywords are more reasonable. At this time, this embodiment will also prompt risks. Only when the two output methods match each other will this embodiment recognize the reliability of the output results.

[0094] S7: Output the operation and maintenance solution text and the retrieval results, and prompt risk information. If a change instruction is received, obtain the changed text input information and / or text description information based on the change instruction, and return to S4. If a reset instruction is received, return to S2.

[0095] Including: using different colors to mark the same parts and different parts in the operation and maintenance solution text and the retrieval results, and covering the display areas of the operation and maintenance solution text and the retrieval results with the prompt of risk information;

[0096] After the confirmation button in the risk information is clicked, clear the risk information and redisplay the operation and maintenance solution text and the retrieval results;

[0097] If the text input information and / or text description information is selected to be changed, display the text input information and / or text description information completely and set it to be editable. After editing based on the change instruction, generate the changed text input information and / or text description information;

[0098] If a reset instruction is received, return to S2.

[0099] On the system interface of this embodiment, the operation and maintenance solution text and the retrieval results are clearly displayed in two adjacent areas. The operation and maintenance solution text is displayed in blue font, and the retrieval results are displayed in green font for the user to quickly distinguish. Among them, the same parts are highlighted with a yellow background, and the different parts are marked with a red background. The prompt of risk information is covered and prompted in the relevant area with a red border and a flashing effect. When the user clicks the confirmation button in the risk information, the system clears the risk prompt and redisplays the operation and maintenance solution text and the retrieval results in normal colors.

[0100] If the user selects to change the text information in the operation and maintenance solution or the retrieval results, the system displays this information completely and sets it to an editable state. The user can directly modify the text content in the edit box, and the system saves the changes in real time and generates the changed text input information and / or text description information.

[0101] At any time, if the user wishes to start over, they can click the "Reset" button on the interface. After receiving the reset instruction, the system clears all current comparison results and edit content and returns to step S2.

[0102] In this embodiment, a multi-modal large model composed of different sub-models is used to support the input and processing of text and images. Among them, the text input is directly converted into text input information, while the image input information and marking information are converted into text description information. The text input information and the text description information are used to generate a collection of text clues to express the complete input content. Subsequently, the present invention separately and independently reads different operation and maintenance plan texts obtained from the text input information and the collection of text clues, and at the same time retrieves historical operation and maintenance plans according to the collection of text clues. Since one of the operation and maintenance plan texts only considers the text input information, while the other operation and maintenance plan text considers both the text input information and the text description information (indirectly considering the image content), when it is uncertain whether the expression of the text or image is clear, the reliability of the two operation and maintenance plan texts is in doubt. Therefore, the present invention continues to verify by comparing historical operation and maintenance plans, and can determine the reliability of the operation and maintenance plan text, and clearly prompt risks while displaying. On the one hand, it can be used as a reference for operation and maintenance personnel, and on the other hand, it does not require operation and maintenance personnel to verify the reliability of each result, improving the operation and maintenance efficiency.

[0103] Embodiment 2: A power grid dispatching automation intelligent operation and maintenance system based on a multi-modal large model, including an operation and maintenance server and an operation and maintenance terminal. The operation and maintenance server and the operation and maintenance terminal are configured to execute the above-mentioned power grid dispatching automation intelligent operation and maintenance method based on a multi-modal large model.

[0104] Among them, the operation and maintenance server mainly executes the deployment and operation of the multi-modal large model to generate an operation and maintenance plan. At the same time, it is responsible for data storage and management, securely stores all operation data, and provides data query and backup functions.

[0105] The operation and maintenance terminal mainly provides a user interface, displays the system status and operation and maintenance plan, and provides input modules such as touch screens, keyboards, microphones, etc., and sends requests to the operation and maintenance server, receives and processes responses.

[0106] The operation and maintenance server and the operation and maintenance terminal work together to achieve the automated and intelligent operation and maintenance of the system, improving the system stability and operation efficiency.

[0107] Embodiment 3: This embodiment also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the above-mentioned power grid dispatching automation intelligent operation and maintenance method based on a multi-modal large model are implemented.

[0108] Embodiment 4: This embodiment also provides a storage medium. Computer-executable instructions are stored in the storage medium, and when the computer-executable instructions are loaded and executed by a processor, the steps of the above-mentioned power grid dispatching automation intelligent operation and maintenance method based on a multi-modal large model are implemented.

[0109] Through the above embodiments, the achievable substantial effects include:

[0110] By constructing a multi-modal large model composed of a language processing model, an image recognition model, and an AIGC operation and maintenance model, the comprehensive processing ability of text and image information is realized. This multi-modal input method enables operation and maintenance personnel to input information in diverse ways, improving the flexibility and accuracy of information input. At the same time, through the image recognition model to perform target detection on the image input information, the key information in the image is converted into text description information, further enriching the description of the operation and maintenance scenario and providing a more comprehensive information basis for the subsequent generation of operation and maintenance plans.

[0111] By separately and independently reading the text input information and the collection of text clues synthesized from the text input information and the text description information, two different operation and maintenance plan texts are output. At the same time, in combination with the retrieval results of historical operation and maintenance plans based on the keywords in the collection of text clues, through the matching degree judgment mechanism, the consistency between the operation and maintenance plan text and the retrieval results is evaluated, which can further verify the reliability and feasibility of the operation and maintenance plan. When the matching degree is lower than the threshold, the system will prompt risk information and allow operation and maintenance personnel to make changes or reset operations as needed. This solves the problem that generative AI in the existing technology may "talk nonsense", thereby improving the accuracy of operation and maintenance decisions.

[0112] By using different colors to mark the same and different parts in the operation and maintenance plan text and the retrieval results, and by covering the display area with the prompt of risk information, etc., the readability and comprehensibility of the operation and maintenance plan are enhanced. These user-friendly designs enable operation and maintenance personnel to grasp key information more quickly, improving the operation and maintenance efficiency.

[0113] In summary, through innovations such as multi-modal information input, dual operation and maintenance plan output, matching degree judgment, and user-friendly display methods, this embodiment realizes a comprehensive understanding and efficient processing of the operation and maintenance scenario, improves the operation and maintenance efficiency and accuracy, and provides a strong technical guarantee for the safe and stable operation of the power system.

[0114] From the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of a specific device is divided into different functional modules to complete all or part of the functions described above.

[0115] In the embodiments provided in the present application, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the embodiments of the structures described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another structure, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of structures or units can be in electrical, mechanical or other forms.

[0116] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0117] In addition, each functional unit in the embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0118] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.

[0119] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for intelligent operation and maintenance of power grid dispatching automation based on a multi-modal large model, characterized in that: The following steps are involved: S1: Pre-build a multimodal large model consisting of a language processing model, an image recognition model, and an AIGC operation and maintenance model; S2: inputting text input information and image input information, and superimposing marking information on the image input information; the marking information is a closed figure drawn by the operator in the image input information; S3: The image recognition model in the multimodal large model recognizes the image input information and the marking information, and outputs text description information; S3 includes: The image recognition model in the multimodal large model performs target detection on the image input information, identifies the scene in the image, outputs scene description information, and performs separate target detection on the image area involved in the marking information, identifies and outputs device description information about the target device in the image area; Based on the preset description template, the scene description information and the device description information are integrated to output text description information; S4: The language processing model in the multimodal large model compares and synthesizes the text input information and the text description information to generate a text clue collection; S4 includes: The language processing model in the multimodal large model recognizes the text input information and text description information, obtains keywords related to the operation and maintenance task as clue information, and represents the extracted clue information in a structured manner; Align and merge the clue information to obtain a collection of text clues; S5: The AIGC operation and maintenance model in the multimodal large model independently reads the text input information and the text clue collection, outputs the corresponding different operation and maintenance solution texts, and searches for historical operation and maintenance solutions based on the keywords of the text clue collection to obtain search results; wherein, the first operation and maintenance solution text is output based on the text input information, and the second operation and maintenance solution text is output based on the text clue collection; S6: Determine the matching degree between different operation and maintenance solution texts and search results. If the matching degrees are all lower than the threshold, execute S7. Otherwise, output the operation and maintenance solution texts and search results whose matching degrees are not lower than the threshold. S7: Output the operation and maintenance plan text and search results, and prompt risk information. If a change instruction is received, obtain the changed text input information and / or text description information based on the change instruction, and return to S4. If a reset instruction is received, return to S2.

2. The method for automatic intelligent operation and maintenance of power grid dispatching based on multimodal large model according to claim 1 is characterized in that: S1: pre-build a multi-modal large model consisting of a language processing model, an image recognition model, and an AIGC operation and maintenance model, including: Building a BERT model based on a number of Transformer encoders, and training the BERT model to obtain a language processing model; Building a YOLO model based on a real-time target detection algorithm, and training the YOLO model to obtain an image recognition model; Based on the Transformer architecture, a GPT model is constructed using a multi-layer self-attention mechanism and position encoding, and the GPT model is trained to obtain the AIGC operation and maintenance model.

3. The method for automatic intelligent operation and maintenance of power grid dispatching based on multimodal large model according to claim 1 is characterized in that: The step S2: inputting text input information and image input information, and superimposing marking information on the image input information, includes: Enter text input information by voice or text input; Input image input information by shooting or importing; The image input information is displayed, an operator draws a closed figure in the image input information to form marking information, the coordinates of the marking information are recorded, and the marking information is superimposed on the image input information based on the coordinate position.

4. The method for automatic intelligent operation and maintenance of power grid dispatching based on multi-modal large model according to claim 1 is characterized in that: S5: The AIGC operation and maintenance model in the multimodal large model independently reads the text input information and the text clue collection, outputs the corresponding different operation and maintenance solution texts, and searches for historical operation and maintenance solutions based on the keywords of the text clue collection to obtain search results, including: The AIGC operation and maintenance model in the multimodal large model reads the text input information and outputs the first operation and maintenance solution text based on the text input information; The AIGC operation and maintenance model in the multimodal large model reads the text clue collection, and outputs the second operation and maintenance solution text based on the text clue collection; At the same time, historical operation and maintenance plans are retrieved, and the keywords in the text clue collection are used to search among the historical operation and maintenance plans, and the operation and maintenance plan corresponding to the training set with the highest relevance is used as the search result.

5. The method for automatic intelligent operation and maintenance of power grid dispatching based on multi-modal large model according to claim 4 is characterized in that: S6: Determine the matching degree between different operation and maintenance solution texts and search results. If the matching degree is lower than the threshold, execute S7. Otherwise, output the operation and maintenance solution text and search results whose matching degree is not lower than the threshold, including: Determine the judgment dimension of the matching degree, compare different operation and maintenance solution texts with the search results based on the judgment dimension, and obtain the matching degree between each operation and maintenance solution text and the search results; If the matching degrees are all lower than the threshold, execute S7; Otherwise, the operation and maintenance solution text and search results whose matching degree is not less than the threshold are output, and the operation and maintenance solution text and search results will be recognized as reliable solutions.

6. The method for automatic intelligent operation and maintenance of power grid dispatching based on multi-modal large model according to claim 1 is characterized in that: S7: outputting the operation and maintenance plan text and search results, and prompting risk information. If a change instruction is received, obtaining the changed text input information and / or text description information based on the change instruction, and returning to S4. If a reset instruction is received, returning to S2, including: Use different colors to mark the same and different parts of the operation and maintenance plan text and search results, and use risk information prompts to cover the display area of ​​the operation and maintenance plan text and search results; After a confirmation button in the risk information is clicked, the risk information is cleared and the operation and maintenance plan text and search results are redisplayed; If the text input information and / or text description information is selected to be changed, the text input information and / or text description information is fully displayed and set to be editable, and after editing based on the change instruction, the changed text input information and / or text description information is generated; If a reset command is received, return to S2.

7. A power grid dispatching automation intelligent operation and maintenance system based on a multi-modal large model, including an operation and maintenance server and an operation and maintenance terminal, characterized in that: The operation and maintenance server and the operation and maintenance terminal are configured to execute the power grid dispatching automation intelligent operation and maintenance method based on a multimodal large model as described in any one of claims 1 to 6.

8. An electronic device, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program, and when the processor calls the computer program in the memory, it implements the steps of the power grid dispatching automation intelligent operation and maintenance method based on a multimodal large model as described in any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium stores computer executable instructions, which, when loaded and executed by a processor, implement the steps of the method for automated intelligent operation and maintenance of power grid dispatching based on a multimodal large model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method of improving image recognition accuracy

    CN106529606A

  • Video generation method, device and system based on multi-modal input

    CN118646940A

  • Power distribution network operation risk visualization method, system, equipment and medium

    CN118940204A

  • Validating answers from an artificial intelligence chatbot

    US20240419988A1

Cited By

  • Intelligent operation and maintenance method and system based on MCP protocol

    CN120614263A