Training method of query converter, pre-training method of multi-modal large model and maintenance method of power system transformer substation
By generating instruction tuning data and introducing cross attention layer into the query converter, the problem of poor compliance with language instructions by multimodal large models is solved, and a query converter and multimodal large models with higher reliability and accuracy are realized, which improves the operating stability and efficiency of the power system.
Patent Information
- Application Number
- CN202510401156.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The existing multimodal large models do not have good compliance with language instructions in the power system, resulting in poor reliability, accuracy and efficiency, mainly because the training data is a picture-text pair rather than a dialogue form, and cannot match the human interaction mode.
By generating instruction tuning data information, including picture content description, several rounds of dialogue and in-depth reasoning, dynamically adjust dialogue rounds, depth and scene adaptability, train the query transformer and splice the cross attention layer after its self-attention layer, establish an interaction mechanism between instructions and query, and improve the correlation of multimodal features.
It improves the reliability and accuracy of the query converter, enhances the task adaptability and scenario understanding capabilities of multimodal large models, and improves the operating stability and efficiency of the power system.
Smart Images

Figure CN120336848A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of large models for power systems, and particularly relates to a training method for a query transformer, a multi-modal large model pre-training method, and a power system substation maintenance method. Background Art
[0002] With the development of economic technology and the improvement of people's living standards, electric energy has become an essential secondary energy source in people's production and life, bringing endless convenience to people's production and life.
[0003] At the present stage, with the advent of the intelligent era, multi-modal large models have been widely applied in power systems, bringing great guarantee to the safe and stable operation of power systems. Therefore, ensuring the stable, reliable, and fast operation of multi-modal large models in power systems is of great significance to power systems.
[0004] Currently, for multi-modal large models in power systems, the MetaLM scheme is mostly adopted, that is, the large language model is used as a unified task layer to connect encoders of different modalities, various tasks are expressed through natural language, and task switching is achieved by means of language instructions. Finally, the generated text is used as the result of the task, and only the loss function of language modeling is used for training. At the same time, for multi-modal large models in power systems at the present stage, Q-Former is also used to complete the extraction and compression of visual features, reduce the number of visual tokens, and keep the visual encoder and the large language model frozen. Only Q-Former with a training parameter quantity in the millions needs to be trained.
[0005] However, at the present stage, there is still a very serious problem with multi-modal large models, that is, although multi-modal large models use language as a unified interface, the compliance effect of multi-modal large models with language instructions is not good enough. This is mainly because the data for training multi-modal models is mainly collected from image-text pairs on the Internet, and these image-text pairs are often not arranged in the form of conversations. Among them, these texts are often descriptions of the content of pictures and cannot match the actual interaction mode between humans and multi-modal large models. This situation makes the reliability, accuracy, and efficiency of current multi-modal large models relatively poor. Summary of the Invention
[0006] One object of the present invention is to provide a training method for a query transformer with high reliability, good accuracy, and high efficiency.
[0007] Another object of the present invention is to provide a multi-modal large model pre-training method including the training method of the query transformer.
[0008] A third object of the present invention is to provide a power system substation maintenance method including the multi-modal large model pre-training method.
[0009] The training method of the query transformer provided by the present invention includes the following steps:
[0010] S1. Obtain picture-text pairs for training the query transformer;
[0011] S2. Based on the picture-text pairs obtained in step S1, generate corresponding instruction tuning data information according to the large language model;
[0012] S3. Input the instruction tuning data information obtained in step S2 into the query transformer to train the query transformer.
[0013] The step of generating corresponding instruction tuning data information according to the large language model based on the picture-text pairs obtained in step S1 in step S2 specifically includes the following steps:
[0014] S2.1. Based on the picture-text pairs obtained in step S1, compile an instruction tuning data generation prompt template; the instruction tuning data generation prompt template includes a description of the picture content, several rounds of conversations about the picture content, and in-depth reasoning about the picture content;
[0015] Among them, the description of the picture content is used to prompt the large language model to generate description data of the picture content; several rounds of conversations about the picture content are used to simulate several rounds of conversations between humans and the model; in-depth reasoning about the picture content is used to control the large language model to design logical reasoning or inference questions based on the picture content;
[0016] S2.2. For each group of picture-text pairs, combine the actual application scenario and semantic requirements, and generate corresponding instruction tuning data information through the large language model to ensure the logicality, coherence, and scenario relevance of the data content;
[0017] During the visual-instruction tuning data generation process, the following parameters are dynamically adjusted:
[0018] Number of dialogue rounds: Control the number of dialogue rounds according to the complexity of the task;
[0019] Depth of dialogue: Control the reasoning depth of the content generated by the model according to the semantic information of the picture-text pair;
[0020] Scene adaptability: Adjust the semantic tendency of the generated content according to business requirements;
[0021] S2.3. Check the semantic consistency of the generated several rounds of dialogue data to ensure the logical rationality, content coherence, and consistency with the scenario requirements of the dialogue context.
[0022] The step of inputting the instruction tuning data information obtained in step S2 in step S3 into the query transformer to train the query transformer specifically includes the following steps:
[0023] Concatenate a cross-attention layer after the self-attention layer of the query transformer;
[0024] Set the instruction input and query interaction mechanism: Input the instructions in the instruction tuning data into the query transformer, and jointly act on the self-attention layer with the learnable queries to complete the cross-modal feature interaction; Establish the association between the queries and the text semantics of the instructions, so that the model can perceive the instruction requirements and adjust the feature extraction direction;
[0025] Instruction-related visual feature extraction: In the cross-attention layer of the query transformer, use the learnable queries after interaction to aggregate visual feature information with a relevance to the instruction higher than the set value from the output of the visual encoder, and ignore the feature noise with a relevance to the instruction lower than the set value to improve the instruction relevance of the multi-modal features;
[0026] Improve the multi-modal task adaptation ability: Train the query transformer with the instruction tuning data information to obtain the trained query transformer.
[0027] The present invention also provides a multi-modal large model pre-training method including the training method of the query transformer, comprising the following steps:
[0028] A. Obtain picture information in the target scenario;
[0029] B. According to the picture information obtained in step A, use a caption generation model to complete the creation of picture-text pairs;
[0030] C. Use the picture-text pairs obtained in step B to train the query transformer trained by the training method of the query transformer to complete the extraction of visual features and the feature alignment between pictures and texts;
[0031] D. Introduce the visual features obtained in step C into the multi-modal large language model to realize the pre-training of the multi-modal large model.
[0032] The step B of using the picture information obtained in step A to complete the creation of picture-text pairs by using a caption generation model specifically includes the following steps:
[0033] B1. Clean the data of the pictures obtained in step A;
[0034] The data cleaning includes deleting pictures that do not meet the set requirements; at the same time, for pictures with a clarity lower than the set requirements, perform image denoising and clarity improvement operations;
[0035] B2. Perform pixel statistics on the pictures after data cleaning;
[0036] Adjust the size of the pictures after data cleaning to a unified size and adjust them to the RGB mode uniformly;
[0037] For the R, G, and B channels of each picture, calculate the pixel mean and pixel variance of each channel;
[0038] According to the pixel mean and pixel variance of each channel, calculate the mean and variance of the R channel, G channel, and B channel respectively;
[0039] B3. Use the caption generation model to generate the description text of the pictures:
[0040] Use the pre-trained caption generation model to generate the corresponding description text for each picture;
[0041] B4. Associate and store the pictures with the corresponding description texts:
[0042] Associate, uniformly number, and name the pictures and the corresponding description files, and store them in the set format.
[0043] The training described in step C specifically includes the following steps:
[0044] C1. Calculate the first loss:
[0045] For the input picture-text pair, calculate the similarity sim from the picture to the text using the following formula i2t [i,j,k]:
[0046] sim i2t [i,j,k] = I i [k]·T j
[0047] In the formula, I i [k] is the feature aggregated by querying similar features on the feature of the i-th picture by the k-th learnable query; T j is the j-th text feature encoded by the query transformer;
[0048] For each picture-text pair, calculate the maximum similarity S from the picture to the text i2t [i,j] is
[0049]
[0050] Calculate the similarity sim from the text to the picture t2i [j,i,k] is sim t2i [j,i,k] = T j ·I i [k];
[0051] For each image-text pair, calculate the maximum similarity from text to image
[0052]
[0053] Finally, calculate the first loss using the following formula ITC :
[0054]
[0055] In the formula, loss i2t is the value of the image-to-text loss function, and loss i2t = CrossEntropy(S i2t [i,j], y); loss t2i is the value of the text-to-image loss function, and loss t2i = CrossEntropy(S t2i [j,i], y); y is the target label; CrossEntropy() is the first intermediate function and n is the total number of categories of the target label;
[0056] C2. Calculate the second loss:
[0057] For the input image-text pair, calculate the similarity sim from image to text i2t [i,j,k] is sim i2t [i,j,k] = I i [k] · T j ;
[0058] For each image-text pair, calculate the maximum similarity S from image to text i2t [i,j] is
[0059]
[0060] For the input image-text pair, calculate the similarity sim from text to image t2i [j,i,k] is sim t2i [j,i,k] = T j · I i [k];
[0061] For each image-text pair, calculate the maximum similarity from text to image
[0062]
[0063] The softmax function is used to convert the similarity matrix into a probability distribution; the diagonal elements of the similarity matrix are modified to a set extremely small value AA to avoid sampling; it is expressed as
[0064] S i2t [i,i] = AA
[0065] S t2i [j,j] = AA
[0066]
[0067] In the formula is the probability that the feature of the i-th image is similar to the feature of the j-th text; is the probability that the feature of the j-th text matches the feature of the i-th image;
[0068] From the obtained probability distribution, the index corresponding to the maximum value is obtained, and the negative sample image-to-text index matrix neg_idx i2t and the negative sample text-to-image index neg_idx t2i ;
[0069] Freeze the image feature I encoded by the visual encoder and the text embedding T; after concatenating the text embedding and the learnable query, input it into the query transformer, and use the image feature as the key-value pair of the cross-attention module of the Transformer encoder in the query transformer, which is expressed as
[0070] output = QFormer(I, Concat(LearnedQueries, T))
[0071] In the formula, output is the output obtained by the learnable query passing through the query transformer; QFormer() is the transformation operation of the query transformer, and the query transformer includes two Transformer encoders; Concat() is the concatenation operation; LearnedQueries is the randomly initialized learnable query;
[0072] Take the first L tensors through the second dimension to obtain the final output output_feature; among them, the dimension of the output of the QFormer() operation is the same as the dimension of Concat(LearnedQueries, T). During implementation, only the first L tensors corresponding to LearnedQueries need to be obtained. Therefore, the first dimension is batch_size, that is, the total number of samples used in a single training iteration; the second dimension is the length of the sequence;
[0073] Through a linearly varying function, the image-text matching score score is obtained as score = Linear(output_feature), where Linear() is the linearly varying function;
[0074] The obtained image-text matching score score is averaged over dimension 1 to obtain the average image-text matching score average_score;
[0075] Finally, the second loss loss is calculated ITM as loss ITM = CrossEntropy(average_score, y);
[0076] C3. Calculate the third loss:
[0077] For the input image-text pair, the following formula is used to calculate the third loss loss IGT :
[0078]
[0079] In the formula, predicted_score is the output of the query transformer, and predicted_score[b, t, v] represents; labels is the probability of the v-th word in the vocabulary at the t-th time step of the b-th sample; labels is the label obtained by converting the text in the image-text pair, and labels[b, t] represents the index of the target word at the t-th time step of the b-th sample; B is the batch size; L is the maximum sequence length;
[0080] C4. In a weighted summation manner, the first loss, the second loss, and the third loss are used as the joint optimization objective and used as the loss function during training to train the model.
[0081] The step of introducing the visual features obtained in step C into the multimodal large language model to perform pre-training on the multimodal large model specifically includes the following steps:
[0082] D1. Introduce a cross-attention module and a gating unit into the multimodal large language model:
[0083] The multimodal large language model includes a number of Transformer blocks; each Transformer block includes a self-attention layer;
[0084] Before the self-attention layer in each Transformer block, add a cross-attention module;
[0085] Meanwhile, between the cross-attention module and the self-attention layer of each Transformer block, a gating unit is added; the gating unit is a tanh function gating unit.
[0086] D2. Language modeling training:
[0087] The language modeling training is the training process after image-text alignment.
[0088] For an image-text pair, the text feature representation after embedding is denoted as T, and the output of the query transformer is denoted as I; the output of the query transformer is used as the key-value pair (K, V) added to the cross-attention module, and cross-attention interaction is performed with the text feature T as the query input, denoted as
[0089]
[0090] In the formula, combined_feature is the mixed feature integrating image-text information; D is the embedding dimension.
[0091] The filtered joint feature output is obtained through the gating unit, denoted as
[0092] gated_feature = tanh(α)·combined_feature
[0093] In the formula, gated_feature is the filtered joint feature output; α is the parameter to be learned.
[0094] The filtered joint feature output passes through the self-attention layer and the feed-forward neural network to obtain the final output, denoted as
[0095] output = ReLU(inter_output·W + b)
[0096] In the formula, output is the final output, and output(b, t, v) represents the probability of the v-th word in the vocabulary at the t-th time step of the b-th sample; ReLU() is the activation function; W is the weight matrix to be learned; b is the bias to be learned; inter_output is the second intermediate variable, and Q is the Q value in the filtered joint feature output gated_feature, K is the K value in the filtered joint feature output gated_feature, and V is the V value in the filtered joint feature output gated_feature.
[0097] During model training, the loss function loss is
[0098]
[0099] The present invention also provides a power system substation maintenance method including the above-mentioned multimodal large model pre-training method, which comprises the following steps:
[0100] (1) Obtain operation and maintenance pictures of substation equipment of the target power system substation;
[0101] (2) Based on the operation and maintenance pictures of substation equipment obtained in step (1), use the above-mentioned multimodal large model pre-training method to train and obtain an operation and maintenance multimodal large model of the target power system substation;
[0102] (3) Use the operation and maintenance multimodal large model of the target power system substation obtained in step (2) to complete the operation and maintenance of the substation equipment of the target power system substation.
[0103] The training method of the query converter, the multimodal large model pre-training method and the power system substation maintenance method provided by the present invention, through the innovative training process of the query converter, and introducing the trained query converter into the multimodal large model for corresponding pre-training, not only realizes the training of the query converter and the pre-training of the multimodal large model, but also the obtained query converter and the pre-trained multimodal large model of the present invention have higher reliability, better accuracy and higher efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] Figure 1 It is a schematic flow chart of the method of the training method of the present invention.
[0105] Figure 2 It is a schematic flow chart of the method of the pre-training method of the present invention.
[0106] Figure 3 It is a schematic flow chart of the method of the power system substation maintenance method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0107] As Figure 1 shown is a schematic flow chart of the method of the training method of the present invention: The training method of the query converter disclosed by the present invention includes the following steps:
[0108] S1. Obtain picture-text pairs for training the query converter;
[0109] S2. Based on the picture-text pairs obtained in step S1, generate corresponding instruction tuning data information according to a large language model (such as the GPT4 model); specifically, it includes the following steps:
[0110] S2.1. Based on the picture-text pairs obtained in step S1, compile an instruction tuning data generation prompt template; the instruction tuning data generation prompt template includes a description of the picture content, several rounds of conversations about the picture content, and in-depth reasoning about the picture content, and the difficulty increases in turn;
[0111] Among them, the picture content description is used to prompt the large language model to generate description data for the picture content. For example, given a picture, the model needs to identify the main elements in it (such as objects, scenes, or text) and generate a concise text description. Example of a prompt template: A picture containing a landscape, please describe the content of the picture / please summarize the content of this picture (using different questioning methods to let the large language model generate a description for the picture).
[0112] Several rounds of conversations about picture content are used to simulate several rounds of conversations between humans and the model. For example, (a picture containing supermarket shelves) You are an AI assistant, please design a conversation between you and a person asking about this photo based on the picture you see. The answer should be in the tone of a visual AI assistant seeing the image and answering questions. Ask different questions and give corresponding answers, including questions about the visual content of the image, which can be about object types, object quantities, object positions, and relative positions between objects, etc. Note: Only ask questions with clear answers.
[0113] Deep reasoning of picture content is used to control the large language model and design logical reasoning or inference questions based on the picture content. For example, you are an AI assistant, please ask complex questions based on the picture content, such as asking about the background knowledge of the objects in the picture, asking about the events happening in the picture, etc. Similarly, do not ask about uncertain details. Provide detailed answers when answering complex questions. For example, give detailed examples or reasoning steps to make the content more persuasive and well-organized. If necessary, you can include multiple paragraphs.
[0114] When generating data, it is necessary to carefully design the images and reasoning tasks to ensure a deep logical connection between the picture content and the questions. Through this template design with gradually increasing difficulty, the model's understanding and processing ability of picture content can be comprehensively improved, from simple description to complex reasoning, to meet the task requirements at different levels.
[0115] S2.2. For each group of picture-text pairs, combined with the actual application scenario and semantic requirements, generate corresponding instruction tuning data information through the large language model to ensure the logicality, coherence, and scenario relevance of the data content.
[0116] By analyzing the semantic features and business requirements of the picture-text pairs, modify the corresponding prompt templates, or also generate prompt templates that combine the actual scenario requirements with the help of the large language model. Based on this, it can be ensured that the generated instruction tuning data can cover specific business scenarios and conform to the corresponding business logic.
[0117] During the generation process of visual-instruction tuning data, dynamically adjust the following parameters:
[0118] Dialogue Turns: Control the number of dialogue turns according to the complexity of the task; for example, for simple description tasks, limit the number of generated turns to within 2 turns; for complex analysis tasks, allow the generation of more than 5 turns of dialogue content;
[0119] Dialogue Depth: Control the inference depth of the content generated by the model according to the semantic information of the image-text pair; such as decomposing the task or inductively analyzing the text information;
[0120] Scene Adaptability: Adjust the semantic tendency of the generated content according to business requirements; for example, generate content related to equipment fault analysis or operation status evaluation in the power industry;
[0121] Finally, based on the dynamically adjusted instruction generation process, ensure that the result of data tuning not only fits the specific business scenario but also achieves a balance among logic, coherence, and practicality;
[0122] S2.3. Perform semantic consistency checks on the generated multiple rounds of dialogue data to ensure the logical rationality, content coherence, and consistency with scene requirements of the dialogue context; a pre-trained semantic matching model (such as Sentence-BERT) can be used to measure the semantic similarity between dialogue turns and filter out dialogues with large semantic deviations; for the filtered dialogue data that does not meet the semantic requirements, it can be regenerated; finally, format and store the generated dialogue data;
[0123] S3. Input the instruction tuning data information obtained in step S2 into the query transformer to train the query transformer; specifically, it includes the following steps:
[0124] Concatenate a cross-attention layer after the self-attention layer of the query transformer;
[0125] Set the instruction input and query interaction mechanism: Input the instructions in the instruction tuning data into the query transformer, and jointly act on the self-attention layer with the learnable queries to complete the cross-modal feature interaction; establish the association between the text semantics of the queries and the instructions, so that the model can perceive the instruction requirements and adjust the feature extraction direction;
[0126] Instruction-related visual feature extraction: In the cross-attention layer of the query transformer, use the interacted learnable queries to aggregate visual feature information with an instruction relevance higher than the set value from the output of the visual encoder, and ignore the feature noise with an instruction relevance lower than the set value to improve the instruction relevance of the multi-modal features; when the Q-Former for instruction tuning is used for multi-modal downstream tasks, through the instruction-driven visual feature extraction ability, it can improve the performance and generalization ability of the multi-modal large model in tasks such as scene understanding, visual question answering, and image-text generation;
[0127] Enhancement of multimodal task adaptation ability: By training a query transformer with instruction tuning data information, a trained query transformer is obtained.
[0128] Such as Figure 2 The following is a schematic flowchart of the pre-training method of the present invention: The pre-training method of the multimodal large model including the training method of the query transformer disclosed in the present invention includes the following steps:
[0129] A. Obtain picture information in the target scenario;
[0130] B. According to the picture information obtained in step A, use a caption generation model to complete the creation of picture-text pairs; specifically, it includes the following steps:
[0131] B1. Clean the data of the pictures obtained in step A;
[0132] The data cleaning includes deleting pictures that do not meet the set requirements, such as picture samples with blurred scenes, unremarkable features, or too high repetition rates; at the same time, for pictures with clarity lower than the set requirements, perform image denoising and clarity improvement operations;
[0133] B2. Perform pixel statistics on the pictures after data cleaning;
[0134] Adjust the pictures after data cleaning to a unified size and uniformly adjust them to the RGB mode;
[0135] For each of the R, G, and B channels of each picture, calculate the pixel mean and pixel variance of each channel;
[0136] According to the pixel mean and pixel variance of each channel, calculate the mean and variance in the R channel, G channel, and B channel respectively; these statistical data will be stored and used for normalization processing in the model training stage; store the mean and variance information for data preprocessing in the subsequent model training and inference processes;
[0137] B3. Use a caption generation model to generate the descriptive text of the pictures:
[0138] Use a pre-trained caption generation model (such as the BLIP model) to generate the corresponding descriptive text for each picture; if the scenario belongs to a professional field (such as the power grid scenario), the general model can be fine-tuned to improve professionalism; conduct manual review on the generated text to ensure correct grammar and close relevance to the image content, and modify inaccurate or incomplete descriptions if necessary;
[0139] B4. Associate and store the pictures with the corresponding descriptive text:
[0140] Associate the pictures with the corresponding description files, assign unified numbers and names, and store them in a set format;
[0141] Specifically, create a main folder (such as "dataset") and set up "images" and "captions" sub-folders respectively; use the os module of Python combined with datetime to generate a unique naming format (such as "YYYYMMDD_HHMMSS_XXX"), and uniformly number and rename the image files and corresponding description texts; store the image files in a unified format (such as.jpg or.png), and store the description texts as files with the same name for easy searching and matching; then, design a JSON structure containing unique identifiers, image paths, and text paths; use the JSON library of Python to generate the corresponding relationship file to ensure clear records and easy reading, providing convenience for model loading or data debugging;
[0142] C. Use the picture-text pairs obtained in step B to train the query transformer trained by the training method of the query transformer to complete the extraction of visual features and the feature alignment of pictures and texts; specifically, it includes the following steps:
[0143] C1. Calculate the first loss:
[0144] For the input picture-text pair, calculate the similarity sim from picture to text using the following formula i2t [i,j,k]:
[0145] sim i2t [i,j,k] = I i [k]·T j
[0146] In the formula, I i [k] is the feature aggregated by the k-th learnable query to query similar features on the i-th picture feature; T j is the j-th text feature encoded by the query transformer;
[0147] For each picture-text pair, calculate the maximum similarity S from picture to text i2t [i,j] is
[0148]
[0149] Calculate the similarity sim from text to picture t2i [j,i,k] is sim t2i [j,i,k] = T j ·I i [k];
[0150] For each image-text pair, calculate the maximum similarity from text to image
[0151]
[0152] Finally, calculate the first loss using the following formula ITC :
[0153]
[0154] where loss i2t is the value of the image-to-text loss function, and loss i2t = CrossEntropy(S i2t [i,j], y); loss t2i is the value of the text-to-image loss function, and loss t2i = CrossEntropy(S t2i [j,i], y); y is the target label; CrossEntropy() is the first intermediate function and n is the total number of categories of the target label;
[0155] C2. Calculate the second loss:
[0156] For the input image-text pair, calculate the similarity sim from image to text i2t [i,j,k] is sim i2t [i,j,k] = I i [k] · T j ;
[0157] For each image-text pair, calculate the maximum similarity S from image to text i2t [i,j] is
[0158]
[0159] For the input image-text pair, calculate the similarity sim from text to image t2i [j,i,k] is sim t2i [j,i,k] = T j · I i [k];
[0160] For each image-text pair, calculate the maximum similarity from text to image
[0161]
[0162] The softmax function is used to convert the similarity matrix into a probability distribution; the diagonal elements of the similarity matrix are modified to a set extremely small value AA to avoid sampling; it is expressed as
[0163] S i2t [i, i] = AA
[0164] S t2i [j, j] = AA
[0165]
[0166] In the formula is the probability that the feature of the i-th image is similar to the feature of the j-th text; is the probability that the feature of the j-th text matches the feature of the i-th image;
[0167] From the obtained probability distribution, the indices corresponding to the maximum values are obtained, and the negative sample image-to-text index matrix neg_idx i2t and the negative sample text-to-image index neg_idx t2i ;
[0168] Freeze the image feature I encoded by the visual encoder and the text embedding T; after concatenating the text embedding and the learnable query, input it into the query transformer, and use the image feature as the key-value pair of the cross-attention module of the Transformer encoder in the query transformer, which is expressed as
[0169] output = QFormer(I, Concat(LearnedQueries, T))
[0170] In the formula, output is the output obtained by the learnable query passing through the query transformer; QFormer() is the transformation operation of the query transformer, and the query transformer includes two Transformer encoders; Concat() is the concatenation operation; LearnedQueries is the randomly initialized learnable query;
[0171] Take the first L tensors from the second dimension to obtain the final output output_feature; among them, the dimension of the output of the QFormer() operation is the same as that of Concat(LearnedQueries, T). During implementation, only the first L tensors corresponding to LearnedQueries need to be obtained. Therefore, the first dimension is the batch size, that is, the total number of samples used in a single training iteration; the second dimension is the length of the sequence;
[0172] Through a linearly varying function, the image-text matching score score is obtained as score = Linear(output_feature), where Linear() is the linearly varying function;
[0173] The obtained image-text matching score score is averaged in dimension 1 to obtain the average image-text matching score average_score;
[0174] Finally, the second loss loss is calculated ITM as loss ITM = CrossEntropy(average_score, y);
[0175] C3. Calculate the third loss:
[0176] For the input image-text pair, the following formula is used to calculate the third loss loss IGT :
[0177]
[0178] In the formula, predicted_score is the output of the query transformer, and predicted_score[b, t, v] represents; labels is the probability of the v-th word in the vocabulary at the t-th time step of the b-th sample; labels is the label converted from the text in the image-text pair, and labels[b, t] represents the index of the target word at the t-th time step of the b-th sample; B is the batch size; L is the maximum sequence length;
[0179] C4. In a weighted summation manner, the first loss, the second loss, and the third loss are used as the joint optimization objective and used as the loss function during training to train the model;
[0180] D. Introduce the visual features obtained in step C into the multi-modal large language model to achieve pre-training of the multi-modal large model; specifically, it includes the following steps:
[0181] D1. Introduce a cross-attention module and a gating unit into the multi-modal large language model:
[0182] The multi-modal large language model includes several Transformer blocks; each Transformer block includes a self-attention layer;
[0183] Before the self-attention layer in each Transformer block, add a cross-attention module;
[0184] Meanwhile, between the cross-attention module and the self-attention layer of each Transformer block, a gating unit is added; the gating unit is a tanh function gating unit.
[0185] D2. Language modeling training:
[0186] The language modeling training is the training process after image-text alignment.
[0187] For an image-text pair, the text feature representation after embedding is denoted as T, and the output of the query transformer is denoted as I; the output of the query transformer is used as the key-value pair (K, V) added to the cross-attention module, and the text feature T used as the query input undergoes cross-attention interaction, denoted as
[0188]
[0189] In the formula, combined_feature is the hybrid feature that fuses image-text information; D is the embedding dimension.
[0190] The filtered joint feature output is obtained through the gating unit, denoted as
[0191] gated_feature = tanh(α) · combined_feature
[0192] In the formula, gated_feature is the filtered joint feature output; α is the parameter to be learned, which is essentially the parameter of the tanh function of the gating unit. This parameter is initially set to 0, making the cross-attention module ineffective at the beginning. The input of the large language model is still pure text features, ensuring the stability of training in this way.
[0193] The filtered joint feature output passes through the self-attention layer and the feed-forward neural network to obtain the final output, denoted as
[0194] output = ReLU(inter_output · W + b)
[0195] In the formula, output is the final output, and output(b, t, v) represents the probability of the v-th word in the vocabulary at the t-th time step of the b-th sample; ReLU() is the activation function; W is the weight matrix to be learned; b is the bias to be learned; inter_output is the second intermediate variable, and Q is the Q value (Query) in the filtered joint feature output gated_feature, K is the K value (Key) in the filtered joint feature output gated_feature, and V is the V value (Value) in the filtered joint feature output gated_feature.
[0196] During model training, the loss function loss is
[0197]
[0198] The present invention also provides a power system substation maintenance method including the above-mentioned multi-modal large model pre-training method, which comprises the following steps:
[0199] (1) Obtain the operation and maintenance pictures of the substation equipment of the target power system substation;
[0200] (2) Based on the operation and maintenance pictures of the substation equipment obtained in step (1), adopt the above-mentioned multi-modal large model pre-training method to train and obtain the operation and maintenance multi-modal large model of the target power system substation;
[0201] (3) Adopt the operation and maintenance multi-modal large model of the target power system substation obtained in step (2) to complete the operation and maintenance of the substation equipment of the target power system substation.
[0202] During specific implementation, first obtain the operation and maintenance pictures of the substation equipment;
[0203] Then, conduct text description manually, including the coordinates of the equipment, the type of the equipment and the status of the equipment; then input the operation and maintenance pictures of the substation equipment and the manually constructed text description into the subtitle generation model to obtain an overall description of this picture, such as "This picture is an operation and maintenance map of substation equipment, in which the vibration damping hammer is in abnormal state and needs to be maintained, and its coordinates in the picture are: (10, 20, 80, 100).".
[0204] Input the constructed substation substation equipment operation and maintenance picture-text pair into large models such as GPT4 to construct an instruction fine-tuning data set; specifically, first construct an instruction tuning data generation prompt template, including picture content description, several rounds of conversations about the picture content and in-depth reasoning about the picture content, etc.; combine the picture-text pair with the prompt template and input it into the large model to obtain the instruction fine-tuning data;
[0205] Pre-train the operation and maintenance multi-modal large model of the target power system substation according to the obtained instruction fine-tuning data set; this large model has powerful picture and text understanding ability and instruction following ability.
[0206] Finally, input the inspection pictures of the substation into the obtained multi-modal large model, then you can ask questions about the content of the pictures, and multi-round conversations and a certain degree of reasoning are supported. For example, ask the multi-modal large model: "What equipment is there in the picture? Does it need to be maintained? Why does it need to be maintained?" etc., so as to realize the maintenance of the power system substation.
Claims
1. A training method for a query transformer, comprising the following steps: S1. Obtain image-text pairs for training the query transformer; S2. Based on the image-text pairs obtained in step S1, generate corresponding instruction tuning data information according to a large language model; S3. Input the instruction tuning data information obtained in step S2 into the query transformer to implement the training of the query transformer.
2. The training method of the query transformer according to claim 1, wherein The step of generating corresponding instruction tuning data information according to the large language model based on the image-text pairs obtained in step S1 in step S2 specifically includes the following steps: S2.
1. Based on the image-text pairs obtained in step S1, write an instruction tuning data generation prompt template; the instruction tuning data generation prompt template includes a description of the image content, several rounds of conversations about the image content, and in-depth reasoning about the image content; Among them, the description of the image content is used to prompt the large language model to generate description data of the image content; several rounds of conversations about the image content are used to simulate several rounds of conversations between humans and the model; in-depth reasoning about the image content is used to control the large language model to design logical reasoning or inference questions according to the image content; S2.
2. For each group of image-text pairs, combine the actual application scenario and semantic requirements, and generate corresponding instruction tuning data information through the large language model to ensure the logicality, coherence, and scenario relevance of the data content; During the visual-instruction tuning data generation process, the following parameters are dynamically adjusted: Number of conversation rounds: Control the number of conversation rounds according to the complexity of the task; Conversation depth: Control the reasoning depth of the content generated by the model according to the semantic information of the image-text pair; Scenario adaptability: Adjust the semantic tendency of the generated content according to business requirements; S2.
3. Perform semantic consistency checks on the generated several rounds of conversation data to ensure the logical rationality, content coherence, and consistency with scenario requirements of the conversation context.
3. The training method of the query transformer according to claim 2, wherein The step of inputting the instruction tuning data information obtained in step S2 in step S3 into the query transformer to implement the training of the query transformer specifically includes the following steps: Concatenate a cross-attention layer after the self-attention layer of the query transformer; Set an instruction input and query interaction mechanism: Input the instructions in the instruction tuning data into the query transformer, and jointly act on the self-attention layer with the learnable queries to complete the cross-modal feature interaction; Establish the association between the text semantics of the queries and the instructions, so that the model can perceive the instruction requirements and adjust the feature extraction direction; Instruction-related visual feature extraction: In the cross-attention layer of the query transformer, use the learnable queries after interaction to aggregate visual feature information with an instruction relevance higher than the set value from the output of the visual encoder, and ignore the feature noise with an instruction relevance lower than the set value to improve the instruction relevance of the multi-modal features; Improve the multi-modal task adaptation ability: Train the query transformer through the instruction tuning data information to obtain the trained query transformer.
4. A multi-modal large model pre-training method including the training method of the query transformer according to any one of claims 1 to 3, characterized in that Including the following steps: A. Obtain image information in the target scenario; B. According to the image information obtained in step A, use a caption generation model to complete the creation of image-text pairs; C. Using the image-text pairs obtained in step B, train the query transformer trained by the training method of the query transformer to complete the extraction of visual features and the feature alignment between images and texts; D. Introduce the visual features obtained in step C into the multi-modal large language model to achieve pre-training of the multi-modal large model.
5. The multimodal large model pre-training method according to claim 4, characterized in that The step B of creating the image-text pairs by using the subtitle generation model according to the image information obtained in step A specifically includes the following steps: B1. Clean the data of the images obtained in step A; The data cleaning includes deleting the images that do not meet the set requirements; at the same time, for the images with clarity lower than the set requirements, perform image denoising and clarity improvement operations; B2. Perform pixel statistics on the images after data cleaning; Adjust the images after data cleaning to a unified size and uniformly adjust them to the RGB mode; For each of the R, G, and B channels of each image, calculate the pixel mean and pixel variance of each channel; According to the pixel mean and pixel variance of each channel, calculate the mean and variance of the R channel, G channel, and B channel respectively; B3. Use the subtitle generation model to generate the description text of the images: Use the pre-trained subtitle generation model to generate the corresponding description text for each image; B4. Associate and store the images with the corresponding description texts: Associate, uniformly number, and name the images and the corresponding description files, and store them in a set format.
6. The multimodal large model pre-training method according to claim 5, wherein The training described in step C specifically includes the following steps: C1. Calculate the first loss: For the input image-text pair, the similarity sim from the image to the text is calculated using the following formula i2t [i, j, k]: sim i2t [i, j, k] = I i [k]·T j where I i [k] is the feature aggregated by querying similar features of the k-th learnable query on the features of the i-th image; T j The j-th text feature encoded for the query transformer; For each image-text pair, calculate the maximum similarity S from the image to the text i2t [i, j] is Calculate the similarity sim between text and image t2i [j, i, k] is sim t2i [j, i, k] = T j ·I i [k]; For each image-text pair, calculate the maximum similarity from text to image Finally, the first loss is calculated using the following formula ITC : where loss i2t is the image-to-text loss function value, and loss i2t = CrossEntropy(S i2t [i,j], y); loss t2i is the text-to-image loss function value, and loss t2i = CrossEntropy(S t2i [j,i], y); y is the target label; CrossEntropy() is the first intermediate function and n is the total number of categories of the target labels; C2. Calculate the second loss: For the input image-text pair, calculate the similarity sim from the image to the text i2t [i, j, k] is sim i2t [i, j, k] = I i [k] · T j ; For each image-text pair, calculate the maximum similarity S from the image to the text i2t [i, j] is For the input image-text pair, calculate the similarity sim from text to image t2i [j, i, k] is sim t2i [j, i, k] = T j ·I i [k]; For each image-text pair, calculate the maximum similarity from text to image Use the softmax function to convert the similarity matrix into a probability distribution; modify the diagonal elements of the similarity matrix to a set minimum value AA to avoid sampling; expressed as S i2t [i,i] = AA S t2i [j,j] = AA where is the probability that the feature of the i-th image is similar to the feature of the j-th text; is the probability that the feature of the j-th text matches the feature of the i-th image; From the obtained probability distribution, obtain the index corresponding to the maximum value, and respectively obtain the negative sample image-to-text index matrix `neg_idx` i2t and the negative sample text-to-image index `neg_idx` t2i ; Freeze the image features I encoded by the visual encoder and the text embeddings T; after concatenating the text embeddings and the learnable queries, input them into the query transformer, and use the image features as the key-value pairs of the cross-attention module of the Transformer encoder in the query transformer, expressed as output = QFormer(I, Concat(LearnedQueries, T)) In the formula, output is the output obtained by the learnable queries passing through the query transformer; QFormer() is the transformation operation of the query transformer, and the query transformer includes two Transformer encoders; Concat() is the concatenation operation; LearnedQueries is the randomly initialized learnable query; Take the first L tensors through the second dimension to obtain the final output output_feature; Obtain the image-text matching score score through the linear transformation function as score = Linear(output_feature), where Linear() is the linear transformation function; Take the average of the obtained image-text matching score score in dimension 1 to obtain the average image-text matching score average_score; Finally, the second loss is calculated as loss ITM where loss ITM = CrossEntropy(average_score, y); C3. Calculate the third loss: For the input image-text pair, the following formula is used to calculate the third loss loss IGT : Where predicted_score is the output of the query transformer, and predicted_score[b, t, v] represents; labels are the probabilities of the v-th word in the vocabulary at the t-th time step of the b-th sample; labels are the labels converted from the text in the image-text pair, and labels[b, t] represents the index of the target word at the t-th time step of the b-th sample; B is the batch size; L is the maximum sequence length; C4. In a weighted summation manner, the first loss, the second loss, and the third loss are used as the joint optimization objective and the loss function during training to train the model.
7. The multimodal large model pre-training method according to claim 6, wherein The step of introducing the visual features obtained in step C into the multimodal large language model to perform pre-training on the multimodal large model specifically includes the following steps: D1. Introduce a cross-attention module and a gating unit into the multimodal large language model: The multimodal large language model includes several Transformer blocks; each Transformer block includes a self-attention layer; Before the self-attention layer in each Transformer block, add a cross-attention module; At the same time, between the cross-attention module and the self-attention layer of each Transformer block, add a gating unit; the gating unit is a tanh function gating unit; D2. Language modeling training: The language modeling training is the training process after image-text alignment; For the image-text pair, the text feature representation after embedding is T, and the output of the query transformer is I; the output of the query transformer is used as the key-value pair (K, V) added to the cross-attention module to perform cross-attention interaction with the text feature T used as the query input, expressed as Where combined_feature is the mixed feature that fuses image-text information; D is the embedding dimension; Obtain the filtered joint feature output through the gating unit, expressed as gated_feature = tanh(α)·combined_feature Where gated_feature is the filtered joint feature output; α is the parameter to be learned; Pass the filtered joint feature output through the self-attention layer and the feed-forward neural network to obtain the final output, expressed as output = ReLU(inter_output·W + b) Where output is the final output, and output(b, t, v) represents the probability of the v-th word in the vocabulary at the t-th time step of the b-th sample; ReLU() is the activation function; W is the weight matrix to be learned; b is the bias to be learned; inter_output is the second intermediate variable, and Q is the Q value in the filtered joint feature output gated_feature, K is the K value in the filtered joint feature output gated_feature, and V is the V value in the filtered joint feature output gated_feature; During model training, the loss function loss is 8. A power system substation maintenance method including the multi-modal large model pre-training method according to any one of claims 4 to 7, characterized in that It includes the following steps: (1) Obtain the operation and maintenance pictures of the substation equipment in the target power system substation; (2) Based on the operation and maintenance pictures of the substation equipment obtained in step (1), use the multimodal large model pre-training method described above to train the operation and maintenance multimodal large model of the target power system substation; (3) Use the operation and maintenance multi-modal large model of the target power system substation obtained in step (2) to complete the operation and maintenance of the substation equipment in the target power system substation.