Electric power image-text large model fine tuning method and system, electronic equipment and storage medium
By fine-tuning the target positioning, image description and question-and-answer of the power graphic and text model, the problem of low accuracy of the multi-modal model of the graphics and text in the power scenario is solved, and the application capabilities of the model in the power field is improved.
Patent Information
- Application Number
- CN202510470216.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-01
AI Technical Summary
When existing multimodal large-scale graphics and text models are used in power scenarios, there is a problem of low accuracy, especially in deep semantic understanding and complex scenario analysis, false alarm rates and omission rates are high, and there is a lack of fine-tuning data sets for the power field.
Through the cyclic iterative training method, the power graphic and text big model is fine-tuned in target positioning, image description and power image Q&A. Using manually marked power graphic and text samples and data sets, a graphic and text Q&A pair and instruction set are formulated to gradually improve the model's target positioning, image description and question-Agreement capabilities.
The goal positioning, image description and question-and-answer of the power graphic and text model in the power field has been improved, and it can adapt to different downstream tasks and output more accurate and practical results.
Smart Images

Figure CN120409736A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to a method for model fine-tuning, and specifically relates to a method, system, electronic device, and storage medium for fine-tuning a large power graph-text model. Background Art
[0002] With the construction of a new power system, the number of devices and the operation frequency have increased significantly, posing higher requirements for intelligent analysis technologies. The object detection and behavior recognition algorithms widely used in the current power industry mostly rely on a single visual data source, but there are high false alarm rates and missed detection rates in deep semantic understanding and complex scene parsing, limiting their application efficiency. At the same time, the power system generates a large amount of multi-modal data such as images and texts every day, which contains rich information and can provide more comprehensive and accurate support for intelligent analysis. However, existing graph-text multi-modal large models are limited by problems such as image resolution differences, inconsistent target scales, and complex characteristics of power core elements in power scenario applications, resulting in low accuracy. Summary of the Invention
[0003] In view of the technical problem of low accuracy when the current graph-text multi-modal large model is applied in power scenarios, this application provides a method, system, electronic device, and storage medium for fine-tuning a large power graph-text model.
[0004] To achieve the above object, this application adopts the following technical solutions: In the first aspect, this application proposes a method for fine-tuning a large power graph-text model, including: Fine-tuning the large power graph-text model through loop iteration training in at least one of the first to third stages according to the downstream task: The first stage: formulating graph-text question-answer pairs for power graph-text sample pairs based on manual annotation, where the answers to the graph-text question-answer pairs include target boxes for characterizing positions added to each power target in the images of the manually annotated power graph-text samples; fine-tuning the target positioning of the pre-trained large power graph-text model through the graph-text question-answer pairs; The second stage: combining the characteristics of power services, setting the image content description order for the data in the manually annotated power graph-text dataset to form an instruction set including multiple image content description order instructions, and fine-tuning the image description of the pre-trained large power graph-text model through the images and the instruction set in the manually annotated power graph-text dataset; The third stage: fine-tuning the power image question-answer of the pre-trained large power graph-text model through the manually annotated power graph-text question-answer pairs.
[0005] In the second aspect, this application proposes a system for fine-tuning a large power graph-text model, including: A stage selection module, configured to determine, according to a downstream task, to fine-tune a large-scale power graph-text model through loop iteration training in at least one of stage one to stage three; A stage one module, configured to formulate graph-text question-answer pairs based on manually annotated power graph-text sample pairs, where the answers to the graph-text question-answer pairs include target boxes for characterizing positions added to each power target in the images of the manually annotated power graph-text samples; and fine-tune the pre-trained large-scale power graph-text model for target localization through the graph-text question-answer pairs; A stage two module, configured to set an image content description order for the data in a manually annotated power graph-text dataset in combination with power business characteristics to form an instruction set including a plurality of image content description order instructions, and fine-tune the pre-trained large-scale power graph-text model for image description through the images and the instruction set in the manually annotated power graph-text dataset; A stage three module, configured to fine-tune the pre-trained large-scale power graph-text model for power image question answering through manually annotated power graph-text question-answer pairs.
[0006] In a third aspect, the present application provides an electronic device, including: a memory, and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned method for fine-tuning a large-scale power graph-text model.
[0007] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned method for fine-tuning a large-scale power graph-text model are implemented.
[0008] Compared with the prior art, the present application has the following beneficial effects: The present application provides a method for fine-tuning a large-scale power graph-text model. In stage one, the pre-trained large-scale power graph-text model is fine-tuned for target localization through graph-text question-answer pairs. In stage two, the pre-trained large-scale power graph-text model is fine-tuned for image description through the images and corresponding instruction sets in a manually annotated power graph-text dataset. In stage three, the pre-trained large-scale power graph-text model is fine-tuned for power image question answering through manually annotated power graph-text question-answer pairs. In view of the lack of a fine-tuning dataset for a large-scale graph-text model in the current power field, the present application formulates corresponding fine-tuning datasets for different downstream tasks according to manually annotated graph-text sample pairs. Through fine-tuning in three stages, the target localization, image description, and question-answering capabilities of the large-scale graph-text model in the power field are respectively improved. The fine-tuned model can adapt to different downstream tasks and output more practical and accurate results.
[0009] The present application also provides a fine-tuning system for a power graph-text large model, an electronic device, and a computer storage medium, which have all the advantages of the above-mentioned power graph-text large model fine-tuning method. Description of the Drawings
[0010] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a schematic flowchart of the second embodiment of the power graph-text large model fine-tuning method of the present application; Figure 2 It is a schematic diagram of target location fine-tuning in the embodiments of the present application; Figure 3 It is a schematic diagram of a power graph-text large model fine-tuning system of the present application. Detailed Embodiments
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations.
[0013] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but merely represents the selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0014] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0015] In the description of the embodiments of the present application, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed during use. This is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present application. In addition, terms such as "first", "second", etc. are only used for differential description and should not be construed as indicating or implying relative importance.
[0016] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.
[0017] In the description of the embodiments of the present application, it should also be noted that unless otherwise clearly specified and limited, if terms such as "set", "installed", "connected", "coupled" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0018] With the construction of the new power system, the number of devices and the operation frequency have increased significantly, which not only improves the complexity and management difficulty of the power system, but also puts forward higher requirements for intelligent analysis technology. As a key means to timely discover, evaluate and solve power equipment defects, personnel violation behaviors and environmental hazards, intelligent analysis technology is crucial for ensuring the safe and stable operation of the power system.
[0019] Currently, most of the object detection and behavior recognition algorithms widely used in the power industry rely on a single visual data source. However, in cases where deep semantic understanding and complex scene analysis are required, these data processing methods based on a single modality often result in a high false alarm rate and a high miss rate, limiting their effectiveness in practical applications. In addition, the power system generates a large amount of multi-modal data such as images and texts every day, and these data contain rich information, which can provide more comprehensive and accurate support for the intelligent analysis of the power system. Therefore, it has become particularly urgent to develop a power graph-text large model that can integrate and process multi-modal data.
[0020] The current research on large multimodal models of graphics and text is mainly devoted to basic capabilities such as architecture research, feature alignment, and pre-training. The performance of downstream tasks is limited by the large differences in power image resolution, different target scales, and complex characteristics of power core elements. The accuracy is low, especially in regression tasks such as target positioning. Therefore, this application proposes a method, system, electronic device, and storage medium for fine-tuning large models of power graphics and text based on the characteristics of power scenarios. A test scheme is also formulated to evaluate the performance of the fine-tuned large model of power graphics and text. The fine-tuning method can be further optimized based on the test results to gradually improve the performance of the large model of power graphics and text. The following is a detailed description of this application in conjunction with the embodiments and drawings.
[0021] As an embodiment of the method for fine-tuning the large power graphics model of the present application, the method may include: Based on the downstream tasks, fine-tune the large power graph and text model through at least one of the stages 1 to 3 through cyclic iterative training: Phase 1: Develop image-text question-answer pairs based on manually annotated power image-text sample pairs. The answers to the image-text question-answer pairs include target boxes added to the images of the manually annotated power image-text samples to represent the location of each power target; use the image-text question-answer pairs to fine-tune target positioning in the pre-trained power image-text large model.
[0022] In practical applications, corresponding image-text question-answer pairs can be formulated based on the needs of downstream tasks, such as equipment defect identification, personnel behavior analysis, or environmental hazard assessment. These question-answer pairs can closely revolve around the core content of the power image-text samples to ensure that the model can learn key information related to the task during fine-tuning. In addition, in the images of the power image-text samples, target boxes are added to represent the location of each power target, such as transformers, poles, switchgear, etc. These target boxes not only indicate the location of the power target, but also provide a basis for subsequent masking operations. During the fine-tuning process, the features of non-target boxes in the image are gradually masked, so that the model will gradually ignore the parts of the image that are not related to the power target during training, thereby focusing more on the visual features of the power target, which helps to improve the model's ability to accurately locate the power target.
[0023] Phase 2: Based on the characteristics of the power business, the image content description order is set for the data in the manually annotated power graph and text dataset to form an instruction set including multiple image content description sequence instructions. The image description of the pre-trained power graph and text large model is fine-tuned through the images and instruction set in the manually annotated power graph and text dataset.
[0024] By setting a reasonable order for describing the image content and forming an instruction set, the large-scale power graph-text model can gradually understand each part of the power image in a logical order. The fine-tuning method in the second stage enables the large-scale power graph-text model to capture detailed information in the power image more comprehensively, enhancing its overall ability to grasp the image content. During the fine-tuning process, the large-scale power graph-text model can learn to generate accurate and coherent image descriptions according to the instruction set, which not only helps the model provide useful information in the power image Q&A task but also provides valuable references for the operation and maintenance and fault diagnosis of power equipment.
[0025] Stage 3: Fine-tune the pre-trained large-scale power graph-text model for power image Q&A using manually annotated power graph-text Q&A pairs.
[0026] In Stage 3, a large number of manually annotated power graph-text Q&A pairs are collected and organized. These Q&A pairs can cover multiple aspects of the power field, ensuring that the model can access rich power knowledge and Q&A scenarios during the fine-tuning process. During the fine-tuning process, the model will learn how to generate accurate answers based on the input questions and power images. This process not only improves the model's Q&A ability but also enhances its comprehensive understanding ability of power images and texts.
[0027] It should be noted that the fine-tuning in Stages 1 to 3 is determined according to the downstream tasks. If target localization needs to be focused on in the downstream task, the large-scale power graph-text model's target localization ability can be strengthened through Stage 1; if image description needs to be focused on in the downstream task, the large-scale power graph-text model's image description ability can be strengthened through Stage 2; if power image Q&A needs to be focused on in the downstream task, the large-scale power graph-text model's power image Q&A ability can be strengthened through Stage 3. In practical applications, if the downstream task involves multiple aspects of concern, the three stages can be combined.
[0028] As Figure 1 shown, the flowchart of the second embodiment of the fine-tuning method for the large-scale power graph-text model of this application may include: S101, construct a fine-tuning dataset for the large-scale power graph-text model.
[0029] For the characteristics of specific power downstream business scenarios, preprocess the power graph-text annotation dataset. Specifically, the following methods can be used: First, make a target localization dataset. Based on a large number of manually annotated power graph-text sample pairs, use the pre-trained base of the large-scale power text model to formulate a graph-text Q&A thought chain. Its input is the manually annotated graph-text data, and the model is required to generate Q&A pairs according to the thought chain mode. For example, for image description : "This figure mainly shows a close-up of a typical transmission tower. The weather in the figure....., there is a suspension clamp hanging at the top of the crossarm above the tower, and a bolt fixing the U-shaped hanging ring on the suspension clamp", add the position of each power target (description in the power targets include transmission towers, crossarms, suspension clamps, bolts, etc.) in the figure, that is, the target box [x1, x2, w, h], and <box>Start< / box> end. The answer should include multiple levels. For each target box, a Q&A pair can be generated. The description There are four power targets in the figure from large to small, so four Q&A pairs can be generated: 1. "Is there a transmission tower in the figure?", "Yes, this figure mainly shows a close-up of a typical transmission tower <box> [0,0,100,300]< / box> "; 2. "Is there a? in the figure?", "Yes, this figure mainly shows a close-up of a typical transmission tower <box> [0,0,100,300]< / box> with a crossarm above the tower <box> [0,0,15,25]< / box> "; 3....... At the same time, in order to improve the sample diversity, for each power image, randomly select 3-5 that are not included in the image from the power target library as negative samples for questioning.
[0030] The datasets used in the two fine-tuning stages of stage two and stage three are both manually annotated power text-image datasets.
[0031] It should be noted that by adding position information (i.e., target boxes) to each power target in the power text-image samples, the power text large model can better understand and locate the power facilities in the image. This not only helps to improve the recognition ability of the power text large model for power facilities, but also provides an accurate position reference for the subsequent formulation of the text-image Q&A thinking chain. The generation process of the Q&A pairs reflects the formulation logic of the text-image Q&A thinking chain. By generating a Q&A pair for each target box and sorting the Q&A pairs according to the size of the power targets, a hierarchical and logically clear Q&A system can be formed. This Q&A system helps the power text large model to gradually and deeply understand the power facilities in the image and their relationships, thereby improving the understanding ability and reasoning ability of the power text large model. Finally, by randomly selecting power targets that are not included in the image from the power target library as negative samples for questioning, the power text large model can be forced to continuously learn and optimize its own performance in the process of distinguishing positive and negative samples. This method not only helps to improve the generalization ability of the model, but also enhances the recognition accuracy of the model for power facilities.
[0032] S102, fine-tune the power text-image large model based on progressive positioning.
[0033] Based on the generated progressive power Q&A pairs, the fine-tuning of the power image-text large model can be divided into three fine-tuning stages in this embodiment. In the first and second fine-tuning stages, the image feature extraction and MLP (Multilayer Perceptron) mapping modules of the power image-text large model are mainly fine-tuned to further align the power image information to the text space; in the third stage of the fine-tuning work, to improve the understanding ability of the power image-text large model for power domain knowledge questions, only the large language model branch is opened.
[0034] It should be noted that several key components of the power image-text large model include: (1) Image feature extraction module: The image feature extraction module is responsible for extracting key features from the input power images. These features usually include edges, textures, colors, etc., which help the model understand the power equipment and scenes in the images.
[0035] (2) MLP mapping module: The MLP mapping module maps the image features to the text space to establish the association between images and texts. Through multi-layer non-linear transformations, the MLP mapping module can convert high-dimensional image features into low-dimensional text representations for subsequent processing.
[0036] (3) Large language model branch: The large language model branch is a pre-trained large language model, such as GPT or BERT, etc. It is responsible for processing text data, understanding the semantic information in the texts, and generating corresponding text outputs. In the power image-text large model, the large language model branch is used to process text descriptions and Q&A related to power images.
[0037] For fine-tuning for different downstream tasks, the specific methods can include: (1) Target localization fine-tuning stage.
[0038] Such as Figure 2As shown, it is a schematic diagram of target location fine-tuning in this embodiment. Due to the complex background and semantic levels of power image data, and the varying sizes of power targets, to accurately locate power targets to support downstream tasks such as equipment defect identification, personnel behavior analysis, and environmental hazard detection, first, during fine-tuning, irrelevant image features are gradually masked to enable the power text-image model to focus on the visual and corresponding text features of power targets. Specifically, when calculating the attention scores between image and semantic tokens, the corresponding image patch and its corresponding visual token are found in the image. For subsequent image generation, the focus is mainly on the image patches involved in the target box. Therefore, the tokens corresponding to the image patches not in the target box are set as Mask (mask). The cross-entropy loss is calculated between the text generated by the model and the answers in the question-answer pairs in the dataset, and the loss function is minimized to optimize the model parameters.
[0039] Accurately locating the targets in power images can support downstream tasks such as equipment defect identification, personnel behavior analysis, and environmental hazard detection. In this first stage, by gradually masking irrelevant image features, the model is focused on the visual and text features of power targets. When calculating the attention scores between image and semantic tokens, the tokens corresponding to the image patches not in the target box are set as Mask, ensuring that the model mainly focuses on the image features within the target box. Finally, the model parameters are optimized by minimizing the cross-entropy loss between the text generated by the model and the answers in the question-answer pairs in the dataset, improving the accuracy of the model in target location.
[0040] (2) Image description fine-tuning stage.
[0041] The target location fine-tuning stage lays the foundation for the power text-image model to locate targets in power multi-service scenarios. On this basis, instruction fine-tuning can be carried out to guide the model to understand power images according to the power business thinking chain.
[0042] When performing image description fine-tuning, the input to the model is the image and the instruction set. As an example of training instructions: "Please describe the picture in the following order: 1) General environmental situation, 2) Power scenario, 3) Main power equipment and personnel (from large to small), 4) Equipment defects or personnel behavior". Through such instructions, the model is guided to first focus on the general environmental situation in the image, such as weather, geographical location, etc., then on the power scenario, such as substations, transmission lines, etc., then identify and describe the main power equipment and personnel, in order from large to small, and finally analyze and point out possible equipment defects or abnormal personnel behavior.
[0043] The target positioning fine-tuning phase has already enabled the large-scale power image and text model to accurately locate power targets by gradually masking irrelevant image features, focusing on the visual and corresponding text features of power targets. Building on this foundation, the goal of the image description fine-tuning phase is to further guide the model's understanding of power images according to the power business thinking chain, improving the model's ability to describe power images. This is crucial for the application of the large-scale power image and text model in downstream tasks such as equipment defect identification, personnel behavior analysis, and environmental hazard assessment. Through image description fine-tuning, the model can more accurately understand the complex information in power images and generate valuable descriptions and analyses according to the power business thinking chain, providing strong support for intelligent management in the power industry.
[0044] (3) Power image question answering fine-tuning stage.
[0045] Building on the fine-tuning phases of target location and image description, the model is fine-tuned using manually annotated power image-text question-answer pairs. In practice, these question-answer pairs can be screened by frontline power industry experts to ensure they meet the analysis requirements of various business scenarios, thereby improving model generalization.
[0046] During fine-tuning, the model is trained on power image-text question-and-answer pairs. During training, the model generates predicted answers based on the input question and image. The predicted answers are compared with manually annotated standard answers, and a loss function (such as cross-entropy loss) is calculated. Using a backpropagation algorithm, the model parameters are optimized to gradually bring the predicted answers closer to the standard answers. Through fine-tuning of the power image-based question-and-answer system, the large power image-text model will be able to better serve the power industry, providing strong support for equipment operation and maintenance, fault diagnosis, and safety management.
[0047] The current large-scale image-text model does not focus on specific targets when calculating the attention score between images and text tokens. The large-scale power image-text model has few samples in the fine-tuning stage, making it difficult for the model to effectively learn the correlation between the visual features and text features of the core elements of power. This application uses a progressive analysis method to improve the target positioning, image description and question-answering capabilities of the large-scale image-text model in the power field.
[0048] In practical applications, the three fine-tuning stages can be iterated cyclically, or the capabilities of a certain stage can be fine-tuned specifically based on the test results.
[0049] S103, large model test of power graphics.
[0050] In this embodiment, the test of the large model of power graphics and text is divided into two dimensions, namely subjective question dimension and objective question dimension. In the objective question stage, power professional experts mainly set questions based on the selected image data, including judgment, single choice, multiple choice, etc. For example, for a transmission image , Example of a single-choice question: "1. What kind of power scenario is shown in the figure? A. Power transmission B. Substation C. Power distribution D. Safety supervision". Each image has 3 - 5 questions. Objective scores can better reflect the model's basic image understanding ability than subjective scores. Additionally, to comprehensively test the model's ability to describe the entire power image, VQA (Visual Question Answering), and object detection, this embodiment can further develop a subjective test plan. For subjective questions, there is no need to annotate the test data. Only experts need to compare and analyze the text generated by the model with the image content. Specifically, as an example, subjective questions can be divided into two main parts, namely, the evaluation of basic cognitive ability and the evaluation of application ability: (1) Evaluation of basic cognitive ability: Basic cognitive ability is mainly used to evaluate the model's overall understanding ability of image content and the fluency of text generation. It is specifically divided into the following five aspects: (1) Completeness: Whether the description of core power elements such as equipment and personnel in the figure is complete.
[0051] (2) Fluency: The statement gets full marks if there are no grammar errors. For example, whether it conforms to grammar rules and whether there are spelling mistakes.
[0052] (3) Image - text relevance: Whether the description is relevant to the image content.
[0053] (4) Scene reasoning: Analyze which power scene the picture belongs to, such as power transmission, substation, power distribution, etc.
[0054] (5) Object counting: Examine the model's ability to count core power elements. For a small number of targets, accurate counting is required. When the number of targets is large, an estimate is acceptable. For example, "There are dozens of insulators in the figure".
[0055] (2) Evaluation of application ability: The application ability of the image - text large model is mainly reflected in its ability to accurately detect abnormal phenomena in power images, such as equipment defects, violation behaviors, environmental hazards, etc.
[0056] (1) Defect or violation identification and description: Whether the abnormal phenomena in the image are accurately inferred and detected. Note: It is not necessary to consider whether the position of the abnormal phenomenon detection box is accurate. A description close to the abnormal position is sufficient.
[0057] (2) False alarm of abnormal phenomena: Whether a normal scene is misjudged as abnormal.
[0058] Organize sample annotation and front - line business experts to comprehensively evaluate the power image - text large model's ability from the above - mentioned dimensions based on power business knowledge and image content, and determine the weight of each dimension according to the specific business scenario to make the evaluation result more targeted.
[0059] In view of the problem that the current large-scale power image-text model lacks a unified evaluation standard, this application constructs an evaluation system for large-scale power image-text models, including two scoring dimensions: subjective and objective. After testing, the qualified large-scale power image-text models are deployed and applied in power scenarios. In actual applications, images that are difficult for the model to distinguish can be collected, re-annotated, and then added to the fine-tuning dataset to iteratively update the model.
[0060] Power business scenarios are complex, and the sizes of power targets are inconsistent. In the same power image, there may be large-scale transmission towers, operating personnel, small fittings, and other power core targets with large scale differences. During the pre-training stage of the large-scale power image-text model, alignment training is carried out using a large number of power image-text positive and negative sample pairs. After training, the model has a certain ability to understand power domain images. However, it is very difficult to directly fine-tune the pre-trained image-text model based on the labeled object detection boxes to directly locate power targets of various sizes. Through the fine-tuning of this application, the large-scale power image-text model is enabled to have the ability to reason layer by layer in complex power business scenarios, and locate a power target layer by layer from large to small, for example, from the tower to the suspension clamp, then to the fixed hanging plate, and finally find the bolt in the image. As Figure 3 shown, it is a schematic diagram of a fine-tuning system for a large-scale power image-text model, which may include: A stage selection module, configured to determine, according to the downstream task, to fine-tune the large-scale power image-text model through at least one of the first stage to the third stage by cyclic iterative training; The first stage module is configured to formulate image-text question-answer pairs for the power image-text sample pairs based on manual annotation, and the answers to the image-text question-answer pairs include target boxes for characterizing positions added to each power target in the images of the power image-text sample pairs based on manual annotation; and perform target localization fine-tuning on the pre-trained large-scale power image-text model; The second stage module is configured to, in combination with the characteristics of power services, set the image content description order for the data in the manually annotated power image-text dataset to form an instruction set including multiple image content description order instructions, and perform image description fine-tuning on the pre-trained large-scale power image-text model through the images and the instruction set in the manually annotated power image-text dataset; The third stage module is configured to perform power image question-answer fine-tuning on the pre-trained large-scale power image-text model through the manually annotated power image-text question-answer pairs.
[0061] In some embodiments of the fine-tuning system for the large-scale power image-text model of this application, in the first stage module, the formulated image-text question-answer pairs include multiple levels of question-answer pairs, and one question-answer pair is generated for each encountered target box.
[0062] In some embodiments of the fine-tuning system for the large-scale power image-text model of this application, the image-text question-answer pairs formulated in the first stage module further include multiple negative sample question-answer pairs.
[0063] In some embodiments of the power graphic model fine-tuning system of the present application, it also includes testing the accuracy of the output results of the power graphic model through subjective tests and objective tests respectively.
[0064] In some embodiments of the power graphic model fine-tuning system of the present application, the order of image content description in the stage two module includes environmental overview, power scene, power equipment and personnel from large to small, equipment defects or personnel behaviors.
[0065] It should be noted that in the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of each module is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components may or may not be physically separated. The components shown as modules can be a physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0066] In addition, in each embodiment of the present invention, each module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0067] The embodiment of the present application also provides an electronic device, which may include one or more processors, a memory, and a communication interface.
[0068] Among them, the memory and the communication interface are coupled to the processor. For example, the memory and the communication interface can be coupled together through a bus.
[0069] Among them, the communication interface is used for data transmission with other devices. The memory stores computer program code. The computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the steps of the above-mentioned power graphic model fine-tuning method.
[0070] Among them, the processor can be a processor or a controller. For example, it can be a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in conjunction with the present disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The processor can be used to support the electronic device in executing the method steps provided in the above embodiments.
[0071] Among them, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above buses can be divided into an address bus, a data bus, a control bus, and so on.
[0072] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned power graphic large model fine-tuning method.
[0073] The computer-readable storage medium involved in the present application includes a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium well-known in the technical field.
[0074] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A fine-tuning method for a large-scale power graph model, characterized in that, Including: According to the downstream task, fine-tune the large-scale power graph-text model through cyclic iterative training in at least one of Phase 1 to Phase 3: Phase 1: Formulate graph-text Q&A pairs based on manually annotated power graph-text sample pairs. The answers in the graph-text Q&A pairs include the target boxes for characterizing positions added to each power target in the images of the manually annotated power graph-text samples. Fine-tune the pre-trained large-scale power graph-text model for target localization through the graph-text Q&A pairs; Phase 2: Combining the characteristics of power operations, set the image content description order for the data in the manually annotated power graph-text dataset to form an instruction set including multiple image content description order instructions. Fine-tune the pre-trained large-scale power graph-text model for image description through the images and the instruction set in the manually annotated power graph-text dataset; Phase 3: Fine-tune the pre-trained large-scale power graph-text model for power image Q&A through manually annotated power graph-text Q&A pairs.
2. The fine-tuning method for a large-scale power graphic model according to claim 1, wherein The method for fine-tuning the large-scale power graph-text model through cyclic iterative training in at least one of Phase 1 to Phase 3 according to the downstream task includes: When the downstream task includes target localization, at least fine-tune the large-scale power graph-text model through cyclic iterative training in Phase 1; When the downstream task includes image description, at least fine-tune the large-scale power graph-text model through cyclic iterative training in Phase 2; When the downstream task includes power image Q&A, at least fine-tune the large-scale power graph-text model through cyclic iterative training in Phase 3.
3. The method for fine-tuning a large power graphic model according to claim 1, wherein The graph-text Q&A pairs formulated in Phase 1 include multiple levels of Q&A pairs, and one Q&A pair is generated for each encountered target box.
4. The method for fine-tuning a large-scale power graph model according to claim 3, wherein, The graph-text Q&A pairs formulated in Phase 1 also include multiple negative sample Q&A pairs.
5. A method for fine-tuning a large-scale power graph model according to claim 1, characterized in that, The method for fine-tuning the pre-trained large-scale power graph-text model for target localization through the graph-text Q&A pairs includes: Taking the graph-text Q&A pairs as input, train the large-scale power graph-text model based on the attention mechanism; Among them, the attention mechanism focuses on the attention scores between the image and the semantic tokens. The cross-entropy loss is used as the loss function during training.
6. The method for fine-tuning a large power graphic model according to claim 1, wherein It also includes testing the accuracy of the output results of the large-scale power graph-text model through subjective tests and objective tests respectively.
7. The method for fine-tuning a large power graphic model according to claim 1, wherein, The image content description order in Phase 2 includes environmental overview, power scene, power equipment and personnel from large to small, equipment defects or personnel behaviors.
8. A fine-tuning system for a large-scale power graph model, characterized in that, Including: A phase selection module for determining to fine-tune the large-scale power graph-text model through cyclic iterative training in at least one of Phase 1 to Phase 3 according to the downstream task; A Phase 1 module for formulating graph-text Q&A pairs based on manually annotated power graph-text sample pairs. The answers in the graph-text Q&A pairs include the target boxes for characterizing positions added to each power target in the images of the manually annotated power graph-text samples. Fine-tune the pre-trained large-scale power graph-text model for target localization through the graph-text Q&A pairs; The second-stage module is used to set the image content description order for the data in the manually annotated power graphic and text dataset in combination with the characteristics of power operations, form an instruction set including multiple image content description order instructions, and perform image description fine-tuning on the pre-trained power graphic and text large model through the images and the instruction set in the manually annotated power graphic and text dataset; The third-stage module is used to perform power image question and answer fine-tuning on the pre-trained power graphic and text large model through the manually annotated power graphic and text question and answer pairs.
9. The fine-tuning system for a large power graphic model according to claim 8, characterized in that, In the first-stage module, the formulated graphic and text question and answer pairs include multiple levels of question and answer pairs, and one question and answer pair is generated for each target box encountered.
10. The fine-tuning system for a large-scale power graphic model according to claim 9, wherein The graphic and text question and answer pairs formulated in the first-stage module also include multiple negative sample question and answer pairs.
11. The fine-tuning system for a large-scale power graph model according to claim 8, wherein, It also includes testing the accuracy of the output results of the power graphic and text large model through subjective tests and objective tests respectively.
12. The fine-tuning system for a large-scale power graphic model according to claim 8, characterized in that, The image content description order in the second-stage module includes environmental overview, power scene, power equipment and personnel from large to small, equipment defects or personnel behaviors.
13. An electronic device, characterized in that, Including: A memory and one or more processors; the memory is coupled to the processor; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the steps of the power graphic and text large model fine-tuning method according to any one of claims 1-7.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the power graphic and text large model fine-tuning method according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Construction and training method of electric power vision multi-granularity pre-training large model
CN115240075A
Multi-modal large language model training method, electronic equipment and storage medium
CN117409431A
Electric power defect image detection method based on image-text question-answer multi-modal model
CN117763107A
Multi-modal large model processing method and device, storage medium and program product
CN119314117A
Method for detecting defects of power equipment and related products
CN119557716A
Cited By
Multi-mode reinforced fine-tuning power detection method and system
CN121033847A