Electric power image-text interaction method and system based on multi-modal large model, and related equipment
By introducing professional visual encoders into multimodal large models and performing feature fusion, the problem of insufficient image detail extraction and professional discrimination capabilities in power scenarios is solved, and higher recognition and analysis capabilities and more accurate natural language instruction responses are achieved.
Patent Information
- Application Number
- CN202510177581.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
The existing multimodal model lacks the ability to extract image details and professionally determine image details in power scenarios, resulting in poor accuracy on strong professional problems such as equipment defects and hidden dangers of personnel behavior.
The professional vision encoder is introduced into the multimodal large model, and the power vision encoder that is equipped with the multimodal large language model is modified through the training power vision encoder to obtain the power graphics and text model. This model combines the features output from the power vision encoder with the output features of the general vision adapter through the feature converter module, improving the ability to analyze images in professional fields.
It improves the ability of multimodal large models to identify and analyze images in professional fields, enhances the accuracy of understanding and responding to natural language instructions, and improves the accuracy of identifying equipment defects and hidden dangers in power scenarios.
Smart Images

Figure CN120125972A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a power graph-text interaction method, system and related equipment based on a multi-modal large model. Background Art
[0002] With the development of smart grids, the amount of data in power systems has increased sharply, especially data from various sensors, monitoring devices, inspection devices, and user terminals. This data contains rich information, such as equipment operating status, environmental parameters, power consumption patterns, etc., which is of great significance for improving the intelligent management level of power systems. Although there are already some applications in the power industry that use artificial intelligence for data analysis, most solutions focus on a single type of data (e.g., only images or only text). This approach limits the understanding of complex scenarios, especially in situations where decisions need to be made by integrating multiple information sources. In addition, traditional methods often struggle to handle unstructured data, leading to problems such as low information utilization rate and slow response speed. Applying graph-text large models to the power industry can more comprehensively analyze and understand multi-source heterogeneous data in power systems. For example, by simultaneously parsing the monitoring video stream (images) and related log files (text) in a substation, the location and cause of faults can be predicted more accurately; by simultaneously providing inspection pictures and operation procedures, various potential hazards and defects under new types of equipment and new procedures can be identified more flexibly without relying on a large number of defect sample annotations.
[0003] Taking QwenVL developed by Alibaba Cloud as an example, general-purpose graph-text large models usually include the following components: Large Language Model (LLM) decoder: Used to generate the text answers, target box positions, and categories required by users. For different sizes of QwenVL, the same parameter-scale QwenLLM is used as the basic language model.
[0004] Visual encoder: The main function of the visual encoder is to process and understand image information. QwenVL adopts the ViT-bigG architecture (2.54B parameters) and uses the pre-trained weights ViT-bigG (2.54B) of OpenClip.
[0005] Visual adapter: Used to transform the dense, low-dimensional (usually 768 to 2048 dimensions) visual image patch features into a small number of LLM tokens (usually 4096 to 8192 dimensions) suitable for the large language model decoder.
[0006] The training of such large models usually includes the following stages: Pre-training phase: In this phase, the goal is to align the features of the visual encoder and decoder using a large amount of image-text pair data. In this phase, the parameters of the language model are frozen and only the visual encoder and visual adapter are optimized.
[0007] Multi-task pre-training phase: In this phase, the multimodal large model is trained using higher-quality image and text multi-task data, which comes from open source image and text tasks and some self-built datasets. At the same time, the resolution of the input image is enlarged to 448 to better capture details. In addition, seven different tasks are used to train the model simultaneously to ensure its generalization ability. All parameters of the entire model are unfrozen and participate in the training.
[0008] Instruction fine-tuning phase: The last phase focuses on improving the model's conversational capabilities and instruction compliance. At this point, the visual encoder is frozen, while the visual adapter and LLM decoder continue to train. The training data comes from the self-instruction generation method of the large model, with the goal of enabling the multimodal model to better understand and respond to natural language instructions, and improve the ability to interact in multi-round conversations and complex scenarios.
[0009] Based on the existing training paradigm of large multi-modal models of graphics and text, the training of large multi-modal models of power graphics and text can be completed. The specific steps are as follows: 1. Build a high-quality target-level multi-task annotation dataset of power graphics and text.
[0010] 2. Based on the large multimodal model of graphics and text trained in general fields, fine-tune the power graphics and text data. At this stage, all parameters of the entire model are unfrozen and participate in the training.
[0011] 3. Freeze the visual encoder and use the large model self-instruction generation method to generate multiple types of dialogue and instruction data for the power graphic annotation data, and use these data to fine-tune the model. At this stage, only the visual adapter and LLM decoder are involved in the training.
[0012] The above model and the unfine-tuned open source model were tested on the same test set, and the results were manually evaluated. It was found that the trained model had good recognition ability for power image scenes, the richness of the description of image details generally exceeded the results of manual annotation, and the answers to power scene-related knowledge were relatively accurate, but there were a lot of hallucinations, and the model accuracy was poor for highly professional problems involving equipment defects, personnel behavior hazards, etc. The performance of the fine-tuned model significantly exceeded the open source general-purpose multimodal model, but it still could not meet actual needs. The main reason for this problem is that the visual encoder of the existing multimodal model is insufficient in detail extraction and professional discrimination capabilities for power scene image data, and the model accuracy improved by fine-tuning is limited. Summary of the invention
[0013] The purpose of the present invention is to address the problems in the above-mentioned prior art and provide an electric power graphic and text interaction method, system and related equipment based on a multimodal large model, introduce professional field visual encoders into the multimodal large model, improve the multimodal large model's ability to recognize and analyze professional field images, and enhance the accuracy of understanding and responding to natural language instructions.
[0014] In order to achieve the above object, the present invention has the following technical solutions: In a first aspect, a method for electric power graphic and text interaction based on a multimodal large model is provided, comprising: Collect electricity pictures and general domain pictures to train the pre-built electricity visual encoder; Build a large multimodal language model, and modify the universal visual encoder of the large multimodal language model through the trained power visual encoder to obtain a large power picture and text model; Construct a multi-task annotation dataset of power graphics and text, and fine-tune the obtained large power graphics and text model; Use the fine-tuned large model of electric power graphics and text to build a service to answer input images and questions.
[0015] As a preferred solution, in the step of collecting power pictures and general field pictures and training a pre-established power vision encoder, the power vision encoder adopts an image model based on the Transformer architecture and uses a mask image reconstruction method to train the power vision encoder.
[0016] As a preferred solution, the multimodal large language model adopts a large model trained in a general field, and a feature converter module is attached after the power vision encoder to convert the features output by the power vision encoder into features with the same dimension as the features output by the built-in vision adapter.
[0017] As a preferred solution, the step of modifying the universal visual encoder of the multimodal large language model through the trained power visual encoder to obtain the power graphic model includes: sending the input picture to the universal visual encoder of the multimodal large language model and the trained power visual encoder at the same time, and then transforming the features output by the power visual encoder into features with the same dimension as the output features of the built-in visual adapter through the feature converter module; fusing the output features of the built-in visual encoder through the visual adapter with the output of the power visual encoder through the feature converter module, and participating in the decoding of the large language model decoder LLM as image tag tokens.
[0018] As a preferred solution, the step of fine-tuning the obtained power graphic model includes: Fix the weights of the general vision encoder, set the learning rates of the language decoder weights and the feature transformation module of the power vision encoder respectively. For the power vision encoder, according to the module number where the weights are located, use an exponentially decaying learning rate, and the closer the module is to the input end, the smaller the learning rate. During the fine-tuning process, select text or text-image training data with a preset probability according to the data source, and filter out the targets and related corpora smaller than the set threshold, and then send them into the power text-image large model for training.
[0019] As a preferred solution, the steps of constructing the power text-image multi-task annotation dataset include: Collect normal pictures and pictures with defects and potential hazards related to power scene equipment, personnel and environment, annotate the collected pictures, and the annotation content includes picture description and the position of each target; obtain the Q&A corpus of the power scene; and obtain the text-image multi-task dataset of the general scene, and mix it with the power dataset according to a set ratio.
[0020] As a preferred solution, the steps of fine-tuning the obtained power text-image large model include: Scale the original picture to two different resolutions. The longest side length of the first resolution does not exceed the first preset value, and the longest side length of the second resolution does not exceed the second preset value. Send the pictures of the first resolution and the second resolution into the vision encoder for encoding respectively; only select some regions from the pictures of the second resolution and send them into the power vision encoder for encoding, so that the number of image tokens formed by the encoding does not exceed 2 times that of the pictures of the first resolution; For the target involved in the corpus with a proportion in the image not exceeding the set value, according to the position and size of the target involved in the corpus, randomly select regions from the picture so that the ratio of the number of image tokens of the target involved in the image to the number of image tokens formed by the corresponding region is within the set interval; randomly select regions from the picture so that the number of image tokens formed by the selected regions meets the requirements; splice the image tokens formed by the two resolutions with the text image tokens according to the position of the original picture corresponding to each image token, and send them into the large language model decoder LLM for decoding.
[0021] As a preferred solution, for the step of using the fine-tuned power text-image large model to build a service and answer the input pictures and questions, during inference, for pictures with the maximum side length exceeding the preset value, scale the pictures to the first resolution and the second resolution according to the maximum side length, send them into the vision encoder for processing respectively, and then send the formed image tokens into the large language model decoder LLM for decoding at the same time.
[0022] As a preferred solution, the power vision encoder selects a model fine-tuned on a power target detection dataset or a power semantic segmentation dataset, and takes out the backbone network of the model feature encoder as the power vision encoder of the power graphic model.
[0023] In the second aspect, a power graphic and text interaction system based on a multimodal large model is provided, comprising: The power visual encoder training module is used to collect power pictures and general domain pictures and train the pre-established power visual encoder; The power image and text large model construction module is used to construct a multimodal large language model and modify the universal visual encoder of the multimodal large language model through the trained power visual encoder to obtain the power image and text large model; The power graph and text large model fine-tuning module is used to construct a multi-task annotation dataset for power graphs and texts, and fine-tune the obtained power graph and text large model; The service building module is used to use the fine-tuned large model of electric power graphics and text to build services and answer input images and questions.
[0024] As a preferred solution, the power graphic model building module adds a feature converter module after the power vision encoder to convert the features output by the power vision encoder into features with the same dimension as the features output by the built-in vision adapter.
[0025] As a preferred solution, the power graphic model construction module simultaneously sends the input image to the universal visual encoder of the multimodal large language model and the trained power visual encoder, and then transforms the features output by the power visual encoder into features with the same dimension as the output features of the built-in visual adapter through the feature converter module; the output features of the built-in visual encoder through the visual adapter and the output of the power visual encoder through the feature converter module are fused as image tag tokens to participate in the decoding of the large language model decoder LLM.
[0026] As a preferred solution, the fine-tuning module of the large electric power graphic model fixes the weights of the general visual encoder, sets the weights of the language decoder and the learning rate of the feature converter module of the electric power visual encoder respectively, and uses an exponentially decaying learning rate for the electric power visual encoder according to the module number where the weight is located. The closer the module is to the input end, the smaller the learning rate. During the fine-tuning process, text or graphic training data is selected according to the data source with a preset probability, and the targets and related corpora that are less than the set threshold are filtered out and sent to the large electric power graphic model for training.
[0027] As a preferred solution, the power graphic large model fine-tuning module scales the original picture to two different resolutions. The longest side length of the first resolution does not exceed the first preset value, and the longest side length of the second resolution does not exceed the second preset value. The pictures of the first resolution and the second resolution are respectively sent to the vision encoder for encoding; only some regions are selected from the pictures of the second resolution and sent to the power vision encoder for encoding, so that the number of image tokens formed by encoding does not exceed 2 times that of the pictures of the first resolution; For the corpus where the proportion of the target in the image does not exceed the set value, according to the position and size of the target involved in the corpus, regions are randomly selected from the picture so that the ratio of the number of image tokens involved in the target to the number of image tokens formed by the corresponding region is within the set interval; regions are randomly selected from the picture so that the number of image tokens formed by the selected regions meets the requirements; the image tokens formed by the two resolutions are spliced with the text image tokens according to the positions of the original pictures corresponding to the respective image tokens and sent to the large language model decoder LLM for decoding.
[0028] As a preferred solution, during inference, for pictures with the maximum side length exceeding the preset value, the service building module scales the pictures to the first resolution and the second resolution according to the maximum side length. After being processed by the vision encoder respectively, the formed image tokens are sent to the large language model decoder LLM for decoding simultaneously.
[0029] In a third aspect, an electronic device is provided, including a processor and a memory. The processor is configured to execute a computer program stored in the memory to implement the power graphic interaction method based on the multi-modal large model.
[0030] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the power graphic interaction method based on the multi-modal large model is implemented.
[0031] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects: The power graphic-text interaction method based on a multimodal large model proposes a power graphic-text large model constructed based on the graphic-text large model in the general domain, without restricting the number of parameters, structure, training method, etc. of the base model used. The graphic-text large model in the general domain usually includes three parts: the large language model decoder LLM, the general vision encoder, and the vision adapter. To better extract the picture features of the power scenario, first, a power vision encoder is trained on power pictures. The power graphic-text interaction method based on the multimodal large model of the present invention introduces a vision encoder in the professional domain into the multimodal large model, sends the output features of the power vision encoder into a new vision adapter, aligns and fuses them with the features of the general vision adapter, and then sends them into the large language model decoder LLM, improving the analysis ability of the multimodal large model for images in the professional domain. Based on the current situation that there are many unlabeled picture samples in the professional domain and few high-quality graphic-text pair samples, the present invention uses unlabeled pictures to train a professional model, and integrates the professional encoder with the model by modifying the vision adapter, reducing the requirement for labeled data volume in model integration, and alleviating the damage of model fine-tuning to the multimodal general ability and knowledge, improving the performance ability of the multimodal large model on industry data.
[0032] Furthermore, according to the resolution of the input picture, the present invention scales the picture into multiple different sizes, sends it into the vision encoder, and fuses all the obtained features and then sends them into the vision adapter. For ultra-high resolution pictures (the width or height of the picture is greater than 4096), the features processed by the vision encoder at different resolutions are sent into the large language model decoder LLM for decoding, and some image tokens at high resolutions are discarded during training, realizing the identification of details of ultra-high resolution pictures by the multimodal large model and improving the recognition and analysis ability of the multimodal large model for small targets.
[0033] It can be understood that the beneficial effects of the second to fourth aspects above can refer to the relevant descriptions in the first aspect above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 Flowchart of the power graphic-text interaction method based on the multimodal large model in the embodiment of the present invention; Figure 2 Schematic diagram of the power graphic-text large model structure constructed in Embodiment 1 of the present invention; Figure 3A schematic diagram of the structure of the large electric power graphic model constructed in the second embodiment of the present invention; Figure 4 Schematic diagram of the structure of the electric power graphic and text interaction system based on the multimodal large model in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0037] See also Figure 1 The embodiment of the present invention proposes a method for electric power graphic interaction based on a multimodal large model, which improves the performance of the multimodal large model on industry data by connecting an industry-specific visual encoder with professional recognition capabilities in parallel to the visual encoder part, and mainly includes the following steps: S1. Collect power pictures and general domain pictures and train the pre-established power visual encoder; S2. Build a large multimodal language model, and modify the general visual encoder of the large multimodal language model through the trained electric power visual encoder to obtain a large electric power graphic model; S3. Build a multi-task annotation dataset for power graphics and text, and fine-tune the obtained power graphics and text model; S4. Use the fine-tuned large model of electric power graphics and text to build a service to answer input images and questions.
[0038] Embodiment 1 This embodiment uses the Masked Image Modeling (MIM) method to pre-train the visual model on a mixed dataset of power pictures and general field pictures, and constructs a large power picture and text model based on the Qwen2-VL 7B model (huggingface model ID: Qwen / Qwen2-VL-7B-Instruct).
[0039] The Qwen / Qwen2-VL-7B-Instruct model is an advanced multi-modal AI model. Qwen2-VL-7B-Instruct has achieved breakthrough results in image and video understanding. It can accurately process images of various resolutions and scales and is good at handling long videos. The model supports the input of visual data such as images and videos, as well as text queries, and facilitates the simultaneous processing of multiple inputs to improve work efficiency. It not only supports English and Chinese but also can recognize multiple language texts in images, such as European languages, Japanese, Korean, Arabic, and Vietnamese. It can dynamically map images to visual tokens, process images of different resolutions, and simulate the human processing method. It can implement multi-modal rotary position embedding (M-ROPE): by decomposing the position embedding into 1D, 2D, and 3D formats, representing text, visual, and video data respectively, to optimize multi-modal data processing.
[0040] The power vision encoder adopts the ViT architecture with about 250M parameters. ViT (Vision Transformer) is an image classification model based on the Transformer architecture. It divides an image into a series of small patches and linearly embeds these patches into a vector space, and then processes them as the input sequence of the Transformer. In this way, ViT can utilize the powerful capabilities of the Transformer in the field of natural language processing to process image data. ViT-bigG is a larger version of the ViT model with about 2.54 billion parameters (2.54B, where B represents Billion). These parameters are obtained through long-term training on a large-scale dataset, so they have learned a certain degree of general feature representation. Using these pre-trained weights can accelerate the training process of new tasks and improve the performance of the model.
[0041] The first embodiment of the present invention is a power graphic-text interaction method based on a multi-modal large model, and the specific steps are as follows: S1. Pre-training of the power vision encoder. In this step, power pictures and general domain pictures are first collected. Then, the model is pre-trained using the method of MAE ([1] He K, Chen X, Xie S, et al. Masked Autoencoders Are Scalable Vision Learners[J]. 2021. DOI: 10.48550 / arXiv.2111.06377.). The model adopts an encoder-decoder architecture. The encoder is ViT-Large, and RoPE, SwiGLU, and LayerScale are used in the model structure. Each module contains a 16-head and 64-dimensional self-attention module, with a total of 24 modules; the decoder is 6 8-head and 64-dimensional Transformer modules. The pre-training data is sampled according to a preset probability based on the data source and category and then sent to the model for pre-training.
[0042] RoPE (Rotary Position Embedding) is a position encoding strategy in the Transformer model. Since the Self Attention of the Transformer has permutation invariance, it is necessary to introduce position encoding to enable the model to perceive the position information of each word in the input sequence. The essence of RoPE is to perform a transformation on the Query and Key vectors formed by two tokens, so that the transformed Query and Key carry position information, and further enable the inner product operation of Attention to automatically perceive the relative position information without any change. RoPE does not have additional parameters that need to be adaptively learned by the model, so it is an efficient encoding method. Although RoPE adopts an absolute position encoding strategy, it can achieve the effect of relative position encoding in combination with the Attention inner product attention mechanism of the Transformer.
[0043] SwiGLU (Swish Gated Linear Units) uses Swish as the activation function. It is often used to enhance the performance of the FFN (feed-forward neural network) layer in the Transformer architecture. By introducing a gating mechanism and the Swish activation function, the expressive ability and non-linear characteristics of the FFN layer are improved. The Swish activation function is a differentiable non-linear function everywhere, with smooth transition characteristics, which helps the model to converge better during training.
[0044] LayerScale is a technique for normalizing the outputs between layers by scaling the input tensors. It is commonly used in Transformer models such as Vision Transformer (ViT) to improve the training stability and depth of the model. A learnable diagonal matrix is added to the output of each residual block to scale the input tensors, thereby normalizing the outputs between layers. This helps to avoid the vanishing or exploding gradient problems in deep learning networks and improves the training dynamics of the model. By normalizing the outputs between layers, LayerScale helps the model to maintain a stable gradient flow during training. LayerScale enables the model to train deeper and larger-capacity networks, thus improving the expressive power and generalization ability of the model.
[0045] S2. Construction of the power graph-text large model. In this step, first, the weights and model code of the Qwen / Qwen2-VL-7B-Instruct model are obtained. Then, its visual encoder is modified. For the input image, it is sent to both the built-in visual encoder of Qwen2-VL and the power visual encoder obtained in step S1. A "feature transformation module" is attached behind the power visual encoder to transform the features output by the power visual encoder into features with the same dimensionality as the output features of the built-in visual adapter (the PatchMerger module of the Visual model). Assuming the size of the input image is H*W, and the patch size of both the built-in visual encoder and the power visual encoder is 16, then the output feature dimensionality of the built-in visual adapter is r*s*3840, and the output size of the power visual encoder is h*w*1024, where h = H / 16, w = W / 16, r = h / 2, s = w / 2. The output dimensionality of the built-in visual adapter is kept consistent with the encoding dimensionality of the LLM, both being 3840. Therefore, a Space2depth operation with a stride of 2 and an MLP layer with an input of 4096 and an output of 3840 are used as the "feature transformation module" to align the features output by the power visual encoder with the built-in visual encoder. The second Linear module of the MLP layer of this module is initialized with a Gaussian distribution with a standard deviation of 1e-4 to reduce the interference of the power visual encoder on the output of the normal Qwen model. Finally, the output features of the built-in visual encoder passing through the visual adapter are fused with the output features of the power visual encoder passing through the feature transformation module and used as image tokens to participate in the decoding of the LLM decoder. The structure of the power graph-text large model constructed in the first embodiment of the present invention is as Figure 2 shown.
[0046] S3. Construct a high-quality target-level power graph-text multi-task annotation dataset and fine-tune the structure of the power graph-text large model constructed in step S2. The specific steps are as follows: S31. Construct a high-quality target-level power graphic and text multi-task annotation dataset. Collect normal pictures related to equipment, personnel, and environment in the power scenario, as well as pictures with defects and potential hazards, and annotate them. The annotation information is stored in json format. The annotation content includes picture description, the location of each target, description, etc. In addition, obtain high-quality question-and-answer corpora for the power scenario and store them in json format. Optionally, obtain a graphic and text multi-task dataset for the general scenario and mix it with the power dataset in a certain proportion.
[0047] S32. Use the dataset in step S31 to fine-tune the structure of the power graphic and text large model. During the fine-tuning process, fix the weights of the general visual encoder, use a smaller learning rate for the weights of the language decoder, use a larger learning rate for the "feature transformer module" of the power visual encoder, and use an exponentially decaying learning rate for the power visual encoder module according to the module number where the weight is located. The closer the module is to the input end, the smaller its learning rate. According to the experimental results, the coefficient of exponential decay is related to the model structure, and the best value is between 0.7 and 0.9. During the fine-tuning process, select text or graphic and text training data with a preset probability according to the data source, and filter out too small targets and related corpora before sending them into the model for training.
[0048] S33. Use the model trained in step S32 to build a service to answer the pictures and questions sent by users.
[0049] Embodiment 2 Compared with Embodiment 1, use step S32a to replace step S32, and use step S33a to replace step S33: S32a: During the model fine-tuning process, the original image is scaled to two different resolutions, the longest side length of the first resolution does not exceed the first preset value (determined according to the model and the size of the video memory, usually between 768 and 1440), and the longest side length of the second resolution does not exceed the second preset value (determined according to the size of the data set image and the target size, usually between 1920 and 4096). The image of the first resolution and the image of the second resolution are respectively sent to the visual encoder for encoding. In order to save the training calculation amount and the video memory required for training, only some areas are selected from the image of the second resolution and sent to the power visual encoder for encoding, so that the number of tokens formed by the encoding does not exceed 2 times that of the image of the first resolution. A feasible method for selecting some areas is that for targets involving targets in the corpus that account for no more than 9% of the image, first, according to the position and size of the target involved in the corpus, randomly select areas from the image so that the ratio of the number of tokens involved in the target to the number of tokens formed in the area is in a specified interval (generally set to 0.001 to 0.3); then randomly select areas from the image so that the number of tokens formed in the selected area meets the requirements. The tokens formed by the two resolutions are concatenated with the text tokens according to the positions of the original images corresponding to each token, and then sent to the LLM decoder. The structure of the constructed power image and text model is as follows: Figure 3 shown.
[0050] S33a: Use the large power graphics model trained in step S32a to build a service. During inference, for images whose maximum side length exceeds 2 / 3 of the second preset, scale them to the first resolution and the second resolution according to the maximum side length, send them to the visual encoder for processing, and then send the formed tokens to the LLM decoder for decoding.
[0051] In a possible implementation, the power visual encoder described in step S1 of the embodiment of the present invention can select a model fine-tuned on a power target detection dataset or a power semantic segmentation dataset, and take out the backbone network backbone of the model feature encoder as the power visual encoder of the power graphic model.
[0052] See also Figure 4 Another embodiment of the present invention further proposes a power graphic and text interaction system based on a multi-modal large model, comprising: The power visual encoder training module 101 is used to collect power pictures and general domain pictures and train the pre-established power visual encoder; The electric power image and text large model construction module 102 is used to construct a multimodal large language model, and to modify the universal visual encoder of the multimodal large language model through the trained electric power visual encoder to obtain the electric power image and text large model; The power graphic and text large model fine-tuning module 103 is used to construct a power graphic and text multi-task annotation dataset and fine-tune the obtained power graphic and text large model; The service building module 104 is used to build a service using the fine-tuned power graphic and text large model to answer the input pictures and questions.
[0053] In a possible implementation, the power graphic and text large model construction module 102 attaches a feature transformer module after the power vision encoder to transform the features output by the power vision encoder into features with the same feature dimension as the output features of the built-in vision adapter.
[0054] In a possible implementation, the power graphic and text large model construction module 102 simultaneously sends the input picture into the general vision encoder of the multi-modal large language model and the trained power vision encoder, and then uses the feature transformer module to transform the features output by the power vision encoder into features with the same feature dimension as the output features of the built-in vision adapter; adds the output features of the built-in vision encoder through the vision adapter and the output of the power vision encoder through the feature transformer module as image token tokens to participate in the decoding of the large language model decoder LLM.
[0055] In a possible implementation, the power graphic and text large model fine-tuning module 103 fixes the weights of the general vision encoder, sets the weights of the language decoder and the learning rate of the feature transformer module of the power vision encoder respectively, uses an exponentially decaying learning rate for the power vision encoder according to the module number where the weights are located, and the closer the module is to the input end, the smaller the learning rate. During the fine-tuning process, text or graphic and text training data are selected with a preset probability according to the data source, and the targets and related corpora smaller than the set threshold are filtered out and then sent into the power graphic and text large model for training.
[0056] In a possible implementation, the power graphic and text large model fine-tuning module 103 scales the original picture to two different resolutions. The longest side length of the first resolution does not exceed the first preset value, and the longest side length of the second resolution does not exceed the second preset value. The pictures with the first resolution and the second resolution are respectively sent into the vision encoder for encoding; only some regions are selected from the pictures with the second resolution and sent into the power vision encoder for encoding, so that the number of image token tokens formed by encoding does not exceed 2 times that of the pictures with the first resolution; For the corpus where the proportion of the target in the image does not exceed the set value, according to the position and size of the target involved in the corpus, randomly select regions from the picture so that the ratio of the number of image tokens corresponding to the target involved in the image to the number of image tokens formed by the corresponding region is within the set interval; randomly select regions from the picture so that the number of image tokens formed by the selected regions meets the requirements; splice the image tokens formed by the two resolutions according to the positions of the original pictures corresponding to each image token, and send them to the large language model decoder LLM for decoding.
[0057] In a possible implementation manner, when the service building module 104 is reasoning, for pictures with the maximum side length exceeding 2 / 3 of the second preset, scale the pictures to the first resolution and the second resolution according to the maximum side length. After being processed by the visual encoder respectively, the formed image tokens are sent to the large language model decoder LLM for decoding at the same time.
[0058] Another embodiment of the present invention also proposes an electronic device, including a processor and a memory, and the processor is used to execute a computer program stored in the memory to implement the power graphic-text interaction method based on the multi-modal large model.
[0059] Another embodiment of the present invention also proposes a computer-readable storage medium, and the computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the power graphic-text interaction method based on the multi-modal large model is implemented.
[0060] The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. For the convenience of description, only the parts related to the embodiments of the present invention are shown above. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. This computer-readable storage medium is non-transitory and can be stored in the storage devices formed by various electronic devices, and can implement the execution process recorded in the method of the embodiments of the present invention.
[0061] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0062] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0063] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0064] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for electric power graphic and text interaction based on a multimodal large model, characterized in that: include: Collect electricity pictures and general domain pictures to train the pre-built electricity visual encoder; Build a large multimodal language model, and modify the universal visual encoder of the large multimodal language model through the trained power visual encoder to obtain a large power picture and text model; Construct a multi-task annotation dataset of power graphics and text, and fine-tune the obtained large power graphics and text model; Use the fine-tuned large model of electric power graphics and text to build a service to answer input images and questions.
2. The electric power graphic and text interaction method based on a multimodal large model according to claim 1 is characterized in that: In the step of collecting power pictures and general domain pictures and training the pre-established power vision encoder, the power vision encoder adopts an image model based on the Transformer architecture and uses a mask image reconstruction method to train the power vision encoder.
3. The electric power graphic and text interaction method based on multimodal large model according to claim 1 is characterized in that: The multimodal large language model adopts a large model trained in a general field, and a feature converter module is attached after the power vision encoder to convert the features output by the power vision encoder into features with the same dimension as the features output by the built-in vision adapter.
4. The electric power graphic and text interaction method based on multimodal large model according to claim 3 is characterized in that: The step of modifying the universal visual encoder of the multimodal large language model through the trained electric power visual encoder to obtain the electric power picture and text large model includes: sending the input picture to the universal visual encoder of the multimodal large language model and the trained electric power visual encoder at the same time, and then transforming the features output by the electric power visual encoder into features with the same dimension as the output features of the built-in visual adapter through the feature converter module; fusing the output features of the built-in visual encoder through the visual adapter with the output of the electric power visual encoder through the feature converter module, and participating in the decoding of the large language model decoder LLM as image tag tokens.
5. The electric power graphic and text interaction method based on multimodal large model according to claim 4 is characterized in that: The step of fine-tuning the obtained large electric power graphic model comprises: The weights of the general visual encoder are fixed, and the weights of the language decoder and the learning rate of the feature converter module of the power visual encoder are set respectively. The power visual encoder is numbered according to the module where the weight is located, and an exponentially decaying learning rate is used. The closer the module is to the input end, the smaller the learning rate. During the fine-tuning process, text or graphic training data is selected according to the data source with a preset probability, and the targets and related corpora that are smaller than the set threshold are filtered out and sent to the large power graphic model for training.
6. The electric power graphic and text interaction method based on multimodal large model according to claim 1 is characterized in that: The steps of constructing a multi-task annotation dataset of electric power graphics and texts include: Collect normal pictures and pictures with defects and hidden dangers related to power scene equipment, personnel and environment, and annotate the collected pictures, including picture description and the location of each target; obtain question and answer corpus of power scenes; and obtain multi-task image and text datasets of general scenes, and mix them with the power dataset according to the set ratio.
7. The electric power graphic and text interaction method based on multimodal large model according to claim 1 is characterized in that: The step of fine-tuning the obtained large electric power graphic model comprises: The original image is scaled to two different resolutions, the longest side length of the first resolution does not exceed the first preset value, and the longest side length of the second resolution does not exceed the second preset value. The image of the first resolution and the image of the second resolution are respectively sent to the visual encoder for encoding; only a part of the area of the image of the second resolution is selected and sent to the power visual encoder for encoding, so that the number of image marker tokens formed by the encoding does not exceed twice that of the image of the first resolution; For targets involved in the corpus whose proportion in the image does not exceed the set value, randomly select areas from the image according to the position and size of the targets involved in the corpus, so that the ratio of the number of image marker tokens involved in the targets to the number of image marker tokens formed in the corresponding areas is within the set interval; randomly select areas from the image so that the number of image marker tokens formed in the selected areas meets the requirements; the image marker tokens formed by the two resolutions are spliced with the text image marker tokens according to the position of the original image corresponding to each image marker token, and sent to the large language model decoder LLM for decoding.
8. The electric power graphic and text interaction method based on multimodal large model according to claim 7 is characterized in that: The fine-tuned large-scale electric power graphic model is used to build a service to answer input images and questions. During inference, for images whose maximum side length exceeds a preset value, the images are scaled to a first resolution and a second resolution according to the maximum side length, and are sent to the visual encoder for processing respectively. The formed image tag tokens are then sent to the large language model decoder LLM for decoding at the same time.
9. The electric power graphic and text interaction method based on multimodal large model according to claim 1 is characterized in that: Select the model fine-tuned on the power target detection dataset or the power semantic segmentation dataset, and take out the backbone network of the model feature encoder as the power visual encoder of the power graphic model.
10. An electric power graphic and text interactive system based on a multimodal large model, characterized in that: include: The power visual encoder training module is used to collect power pictures and general domain pictures and train the pre-established power visual encoder; The power image and text large model construction module is used to construct a multimodal large language model and modify the universal visual encoder of the multimodal large language model through the trained power visual encoder to obtain the power image and text large model; The power graph and text large model fine-tuning module is used to construct a multi-task annotation dataset for power graphs and texts, and fine-tune the obtained power graph and text large model; The service building module is used to use the fine-tuned large model of electric power graphics and text to build services and answer input images and questions.
11. The electric power graphic and text interactive system based on multimodal large model according to claim 10, characterized in that: The power graphic large model construction module adds a feature converter module after the power visual encoder, which is used to convert the features output by the power visual encoder into features with the same dimension as the features output by the built-in visual adapter.
12. The electric power graphic and text interactive system based on multimodal large model according to claim 11, characterized in that: The electric power graphic and text large model construction module simultaneously sends the input image to the universal visual encoder of the multimodal large language model and the trained electric power visual encoder, and then transforms the features output by the electric power visual encoder into features with the same dimension as the output features of the built-in visual adapter through the feature converter module; the output features of the built-in visual encoder through the visual adapter and the output of the electric power visual encoder through the feature converter module are fused as image tag tokens to participate in the decoding of the large language model decoder LLM.
13. The electric power graphic and text interactive system based on multimodal large model according to claim 12, characterized in that: The electric power graphic and text large model fine-tuning module fixes the weights of the general visual encoder, sets the weights of the language decoder and the learning rate of the feature converter module of the electric power visual encoder respectively, and uses an exponentially decaying learning rate for the electric power visual encoder according to the module number where the weight is located. The closer the module is to the input end, the smaller the learning rate. During the fine-tuning process, text or graphic training data is selected according to the data source with a preset probability, and the targets and related corpora that are less than the set threshold are filtered out and sent to the electric power graphic and text large model for training.
14. The electric power graphic and text interactive system based on multi-modal large model according to claim 10, characterized in that: The power graphic model fine-tuning module scales the original image to two different resolutions, where the longest side length of the first resolution does not exceed a first preset value, and the longest side length of the second resolution does not exceed a second preset value, and the image of the first resolution and the image of the second resolution are respectively sent to the visual encoder for encoding; Only a part of the area is selected from the second resolution image and sent to the power vision encoder for encoding, so that the number of image tokens formed by encoding does not exceed twice that of the first resolution image; For targets involved in the corpus whose proportion in the image does not exceed the set value, randomly select areas from the image according to the position and size of the targets involved in the corpus, so that the ratio of the number of image marker tokens involved in the targets to the number of image marker tokens formed in the corresponding areas is within the set interval; randomly select areas from the image so that the number of image marker tokens formed in the selected areas meets the requirements; the image marker tokens formed by the two resolutions are spliced with the text image marker tokens according to the position of the original image corresponding to each image marker token, and sent to the large language model decoder LLM for decoding.
15. The electric power graphic and text interactive system based on multi-modal large model according to claim 14, characterized in that: When the service building module is inferring, for images whose maximum side length exceeds a preset value, the images are scaled to a first resolution and a second resolution according to the maximum side length, and are sent to the visual encoder for processing respectively. The formed image tag tokens are then sent to the large language model decoder LLM for decoding at the same time.
16. An electronic device, characterized in that: It comprises a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the electric power graphic-text interaction method based on a multimodal large model as claimed in any one of claims 1 to 9.
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, the power graphic-text interaction method based on the multimodal large model as claimed in any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Method for guiding backbone network to carry out image segmentation based on thousand-ask large model
CN120318522A
Image segmentation method based on Qianwen large model to guide backbone network
CN120318522B
Multi-modal large model training method and device, equipment and storage medium
CN120632470A
Multi-modal feature splicing method and multi-modal data processing method based on rotation position coding technology
CN120724380A
Power transaction data interaction method and system based on large model and agent collaboration
CN120765383A