Power transmission line operation and maintenance algorithm and device based on multi-modal pre-training large model, and computer readable storage medium

Through the transmission line operation and maintenance algorithm based on multimodal pre-trained large model, the problems of low multimodal interaction efficiency and poor accuracy of hidden danger identification in transmission line operation and maintenance are solved, and efficient hidden danger identification and accuracy are achieved.

CN120164079AInactive Publication Date: 2025-06-17SHANDONG ZHIYANG ELECTRIC

Patent Information

Application Number
CN202510234729.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the interaction efficiency between multimodals is low, resulting in low accuracy and reliability of hidden danger identification in power transmission line operation and maintenance.

Method used

The transmission line operation and maintenance algorithm based on multimodal pre-trained large models is adopted. By obtaining hidden danger images and text descriptions in the transmission channel scenario, the large-scale language models Qwen and ViT image processing models are used for feature extraction and mapping fusion, and fine-tuning is combined with Lora Adaptive and Q-Lora technologies to improve the model's prediction capabilities in business scenarios.

Benefits of technology

It improves the accuracy and reliability of hidden danger identification in power transmission line operation and maintenance, realizes efficient interaction and deep fusion between multimodals, and enhances the generalization ability and task adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164079A_ABST
    Figure CN120164079A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing based on a specific computer model, and more specifically relates to a power transmission line operation and maintenance algorithm and device based on a multi-modal pre-training large model, and a computer readable storage medium. The method comprises the following steps: S1, acquiring a hidden danger image in a power transmission channel scene, and attaching text description information to the hidden danger of the image to form an image-text sample format; s2, establishing a multi-modal pre-training large model; s3, carrying out fine training operation on the established multi-modal pre-training large model, wherein the fine training operation comprises carrying out lightweight personalized adaptation by using Lora Adaptive and Q-Lora and designing fine tuning parameters; and S4, training and verifying the multi-modal pre-training large model. According to the method, the problems of low interaction efficiency among multiple modes and low accuracy and reliability of hidden danger identification in operation and maintenance of the power transmission line are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing based on specific computer models, and more specifically, relates to a power transmission line operation and maintenance algorithm, device, and computer-readable storage medium based on a multi-modal pre-trained large model. Background Art

[0002] Power transmission visualization devices can effectively alleviate the contradiction between the demand for power grid construction and the cost of manual operation and maintenance, and have the advantages of low cost, simple operation, and being unaffected by the environment. Among the images of power transmission lines taken by visualization, there are many types of hidden dangers to be inspected, including various hidden danger types such as fireworks, foreign objects on conductors, and construction machinery.

[0003] Chinese invention patent CN 117412147A discloses a video summary generation method based on large model fine-tuning. Video summary generation is to summarize and generalize the original video content using text, and is widely used in the multi-modal field. The video summary generation method based on large model fine-tuning technology proposed by the present invention includes the following steps: (1) Utilize the multi-modal fusion characteristics of the MACAW-LLM large model to perform cross-modal interaction between video features and text features; (2) Use CLIP and WHISPER to complete the extraction of video and audio features; (3) Use GPT-3.5Turbo to generate instructions to assist the summary generation algorithm; (4) Use the attention mechanism algorithm to complete modal alignment and fuse it with the instructions; (5) Use the LoRA fine-tuning technology to train the model, and minimize the negative log-likelihood function to iteratively update the large model parameters and generate a video summary.

[0004] In summary, due to the influence of the appearance frequency and collection difficulty in the prior art, it is difficult to collect samples of some categories, and sufficient samples cannot be provided for the training of deep learning algorithms, and the interaction efficiency between multi-modalities is low; moreover, the power transmission scenario is complex and the category features are variable, so the accuracy and reliability of hidden danger identification in power transmission line operation and maintenance are low. Therefore, power transmission line inspection based on images taken by visualization devices still faces huge challenges. Summary of the Invention

[0005] The present invention aims to overcome at least one defect of the above prior art, and provides a power transmission line operation and maintenance algorithm based on a multi-modal pre-trained large model to solve the problems of low interaction efficiency between multi-modalities and low accuracy and reliability of hidden danger identification in power transmission line operation and maintenance.

[0006] The present invention also discloses a device for a power transmission line operation and maintenance algorithm based on a multi-modal pre-trained large model.

[0007] The detailed technical solution of the present invention is as follows:

[0008] A power transmission line operation and maintenance algorithm based on a multi-modal pre-trained large model, the method comprising:

[0009] S1. Obtain a hidden danger image in a power transmission channel scenario, and attach text description information to the hidden danger in the image to form an image-text sample format;

[0010] S2. Establish a multimodal pre-training large model, including a large-scale language model, a ViT image processing model and a feature mapping module; the large-scale language model and the ViT image processing model process the text data and the image data respectively, and then the feature mapping module maps and features the processed text vector and image vector;

[0011] It adopts the advanced Encoder-Decoder architecture as the core support, and specially selects the large-scale language model Qwen as the cornerstone. The model contains pre-trained weights of up to 7 billion parameters. Furthermore, it integrates the VisionTransformer or ViT technology to deeply refine the image features and realize high-fidelity encoding of visual information. More importantly, the present invention designs a novel feature mapping fusion mechanism, which cleverly integrates the image vector extracted by ViT with the text vector optimized by Qwen, ensuring that the information of the two modes is not only seamlessly connected in space, but also deeply interwoven at the semantic level, thereby forging a highly concentrated and generalized unified multimodal representation framework. This framework not only promotes efficient interaction between modalities, but also provides an unprecedentedly strong foundation for the subsequent execution of multimodal tasks.

[0012] S3. Perform fine-tuning training on the established multimodal pre-trained large model, including lightweight personalized adaptation using Lora Adaptive and Q-Lora and designing fine-tuning parameters;

[0013] S4. Train and verify the multimodal pre-trained large model, and then use it to detect hidden danger images of transmission channels.

[0014] Furthermore, the S1 specifically includes:

[0015] S11, collecting and acquiring hidden danger images in the power transmission channel scenario;

[0016] S12. Manually attach text description information to the acquired image hidden dangers.

[0017] Furthermore, a multimodal pre-training large model is established, including a large-scale language model, a ViT image processing model and a feature mapping module; the large-scale language model and the ViT image processing model process the text data and the image data respectively, and then the feature mapping module maps and features the processed text vector and image vector; S2 specifically includes:

[0018] S21. Application and Optimization of the Large Language Model Qwen: Based on the Encoder-Decoder architecture, the Encoder adopts the encoder structure of the large language model Qwen, and the Decoder adopts the decoder of the large language model Qwen;

[0019] The input of the large language model Qwen is preprocessed text data, and the preprocessing includes word segmentation, tokenization, and conversion into a format recognizable by the model;

[0020] The input text data first enters the Encoder. The Encoder processes the input word vectors or character sequences one by one, captures the lexical relationships, syntactic structures, and semantic information in the text through the feature extraction layer of the Encoder, and generates a high-dimensional encoded feature representation; then, the context information is fused with the current input text data; finally, it is input into the Decoder for decoding to obtain an optimized text vector representation; the feature extraction layer of the Encoder includes multiple self-attention mechanisms and feed-forward neural networks;

[0021] S22. Innovative Integration of the ViT Image Processing Model and Image Feature Extraction. Among them, the encoder of ViT is stacked by multiple Transformer blocks, each Transformer block includes multi-head self-attention and a feed-forward neural network, and residual connections and layer normalization are applied between each Transformer block;

[0022] The input of the ViT image processing model is image data. First, the image is segmented into multiple patches, and then the self-attention mechanism is applied to process these patches to capture the global dependencies and subtle features in the image; then, the feature extraction layer of the encoder is used to deeply analyze the input image, identify the basic components of the image, and precisely capture the subtle details and complex structures in the image to obtain a high-dimensional image vector representation, achieving high-fidelity encoding of visual information and exceeding the performance limitations of traditional image feature extraction techniques;

[0023] After completing the image feature extraction, these high-dimensional feature vectors are prepared for subsequent multimodal fusion to ensure that the image information can be effectively represented in a form suitable for combination with other modal data.

[0024] S23. Design an innovative feature mapping scheme aimed at effectively docking the high-dimensional image vectors extracted from ViT with the text vectors optimized by Qwen. This requires developing a bimodal-compatible representation space that can retain the unique information of each modality and promote the mutual understanding and fusion of the two.

[0025] Input the high-dimensional image vectors obtained by ViT and the text vectors optimized by Qwen into the feature mapping module, and map the feature vectors of the two modalities into a bimodal-compatible representation space; in this space, the image and text features are aligned and corresponding to each other, and semantic similarity calculation and context correlation analysis are introduced;

[0026] First, semantic similarity calculation identifies the similar and related parts semantically between the image and text features by calculating the semantic similarity between them; then, context correlation analysis fuses the context features of the image and text to ensure that they can corroborate and complement each other at the semantic level;

[0027] Stack the features after fusion at the semantic level to generate fused features;

[0028] Adopt the method of weighted summation to convert the fusion result of the image and text features into a unified fused feature representation, and the final output is the multimodal fused feature after feature mapping and fusion processing.

[0029] During the process of feature mapping and fusion, special attention is paid to achieving deep interweaving of the information of the two modalities at the semantic level. By introducing advanced semantic analysis techniques, such as semantic similarity calculation, context correlation analysis, etc., it is ensured that the high-level meanings of the image and text can corroborate and complement each other, forming a richer and more accurate multimodal context understanding.

[0030] Furthermore, the specific steps of S3 are as follows:

[0031] S31. Apply Lora Adaptive and Q-Lora for lightweight personalized adaptation:

[0032] Select a representative small-scale high-quality data set from the power transmission business data;

[0033] Introduce the Lora Adaptive technology, add low-rank matrices to each feature extraction layer of the large multimodal pre-trained model to achieve efficient expansion of parameters without significantly increasing the computational burden; at the same time, adopt the quantized version of the Lora technology, add Lora at the decoder position of the large multimodal model, and compress the model size by reducing the magnitude and bit width of the model parameters;

[0034] Furthermore, reducing the magnitude and bit width of the model parameters is to change from double to int8 type.

[0035] S32. Full-scale fine-tuning to achieve deep business customization:

[0036] Design fine-tuning parameters, and during the full-scale fine-tuning process, adopt regularization, early stopping, and data augmentation strategies to effectively monitor and control the model complexity; the fine-tuning parameters include but are not limited to the layer range, learning rate strategy, and batch size.

[0037] In another aspect of the present invention, there is provided a device for implementing a transmission line operation and maintenance algorithm based on a multi-modal pre-trained large model, the device comprising:

[0038] At least one processor; and

[0039] A memory storing instructions that, when executed by the at least one processor, cause the at least one processor to execute the transmission line operation and maintenance algorithm based on the multi-modal pre-trained large model as described above.

[0040] In another aspect of the present invention, there is also provided a machine-readable storage medium storing executable instructions that, when executed, cause the machine to execute the transmission line operation and maintenance algorithm based on the multi-modal pre-trained large model as described above.

[0041] Compared with the prior art, the beneficial effects of the present invention are:

[0042] The transmission line operation and maintenance algorithm, device, and computer-readable storage medium based on the multi-modal pre-trained large model provided by the present invention include obtaining hidden danger images in the transmission channel scenario, attaching text description information to the image hidden dangers, using Qwen based on the Encoder-Decoder architecture and importing the pre-trained parameters of 7B as the large language model, using VisionTransformer to encode the image features, stacking the encoded image vectors and text vectors after mapping, and performing Lora Adaptive fine-tuning on the multi-modal pre-trained large model using the transmission business data, making the large model tend to the business scenario, retaining the prediction ability in the open scenario while realizing the prediction ability on the business scenario data, and applying it to the detection of hidden dangers in the transmission line operation and maintenance, effectively improving the accuracy and reliability of hidden danger identification in the transmission line operation and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flowchart of the transmission line operation and maintenance algorithm based on the multi-modal pre-trained large model of the present invention.

[0044] Figure 2 is a schematic diagram of the image hidden danger in Embodiment 1 of the present invention. DETAILED DESCRIPTION

[0045] The present invention will be further described below with reference to the drawings and embodiments.

[0046] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0047] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0048] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0049] Embodiment 1

[0050] Refer Figure 1 , this embodiment provides a power transmission line operation and maintenance algorithm based on a multi-modal pre-trained large model, and the method includes:

[0051] S1. Construct a model training data set.

[0052] Obtain hidden danger images in the power transmission channel scenario, and attach text description information to the image hidden dangers to form a sample format of image-text.

[0053] Specifically, the S1 includes:

[0054] S11. Collect and obtain hidden danger images in the power transmission channel scenario;

[0055] S12. Manually attach text description information to the obtained image hidden dangers, such as Figure 2 shown in the figure, there is a foreign object on the conductor, and a black and gray cloth bag is hung above the conductor; the attached text description is: The weather is clear in this image, there are poles and conductors, there is a house below the conductor, and a black and gray cloth bag is hung on the conductor, which is likely to cause a short circuit of the line and trigger a fire, and on-site disposal is required.

[0056] S2. Establish a multi-modal pre-trained large model, including a large-scale language model, a ViT image processing model, and a feature mapping module; the large-scale language model and the ViT image processing model process text data and image data respectively, and then the feature mapping module maps and fuses the processed text vectors and image vectors.

[0057] Select the large language model Qwen as the cornerstone and based on the Encoder-Decoder architecture. This model contains pre-trained weights with up to 7 billion parameters. Further, it integrates the Vision Transformer (ViT) technology to deeply refine image features and achieve high-fidelity encoding of visual information.

[0058] Of particular importance is the design of a novel feature mapping fusion mechanism. This mechanism cleverly integrates the image vectors refined by ViT with the text vectors optimized by Qwen, ensuring that the information of the two modalities is not only seamlessly docked in space but also deeply intertwined at the semantic level, thus forging a highly condensed and generalization-capable unified multi-modal representation framework. This framework not only promotes efficient interaction between modalities but also provides an unprecedentedly powerful foundation for subsequent multi-modal task execution.

[0059] Specifically, S2 includes the following detailed steps:

[0060] S21. Application and optimization of the large language model Qwen. Based on the Encoder-Decoder architecture, the Encoder uses the encoder structure of the large language model Qwen, and the Decoder uses the decoder of the large language model Qwen; utilize the 7 billion parameter weights obtained from the pre-training of the Qwen model. These weights are trained on a large corpus and can deeply understand language structure and context, enhancing the fineness and accuracy of the model in natural language processing tasks.

[0061] The specific processing process of the large language model Qwen is as follows:

[0062] The input of the large language model Qwen is preprocessed text data. The preprocessing includes tokenization, tagging, and conversion into a format recognizable by the model, such as word vectors or character sequences, etc.; when dealing with certain specific tasks, such as dialogue generation or text continuation, the model also needs to receive additional context information.

[0063] The input text data first enters the Encoder. The Encoder processes the input word vectors or character sequences one by one, captures the lexical relationships, syntactic structures, and semantic information in the text through the feature extraction layer of the Encoder, and generates a high-dimensional encoded feature representation. Then, the context information is fused with the current input text data. Finally, it is input into the Decoder for decoding to obtain an optimized text vector representation. The feature extraction layer of the Encoder includes multiple self-attention mechanisms and feed-forward neural networks. The Decoder uses the decoder of the large language model Qwen. The encoded features will be passed to the Decoder part for further processing. The Decoder adopts a structure matching the Encoder and generates corresponding outputs according to the encoded features and task requirements.

[0064] Through the above specific processing process, the large language model Qwen can make full use of the rich parameter weights obtained from its pre-training, deeply understand the language structure and context, and thus achieve improvements in fineness and accuracy in various natural language processing tasks.

[0065] S22, Innovative integration of the ViT (Vision Transformer) image processing model and image feature extraction; among them, the encoder of ViT is stacked by multiple Transformer blocks, each Transformer block includes multi-head self-attention and a feed-forward neural network, and residual connections and layer normalization are applied between each Transformer block;

[0066] The input of the ViT image processing model is image data. First, the image is segmented into multiple patches, and then the self-attention mechanism is applied to process these patches to capture the global dependencies and subtle features in the image. Then, the feature extraction layer of the encoder is used to deeply analyze the input image, identify the basic components of the image, and precisely capture the subtle details and complex structures in the image, obtaining a high-dimensional image vector representation, achieving high-fidelity encoding of visual information and exceeding the performance limitations of traditional image feature extraction techniques;

[0067] After completing the image feature extraction, these high-dimensional feature vectors are ready for subsequent multi-modal fusion, ensuring that the image information can be effectively represented in a form suitable for combination with other modal data.

[0068] Using the self-attention layer of ViT, the input image is deeply analyzed, not only identifying the basic constituent elements of the image, but also precisely capturing the subtle details and complex structures in the image, achieving high-fidelity encoding of visual information and surpassing the performance limitations of traditional image feature extraction techniques. After the image feature extraction is completed, these high-dimensional feature vectors are ready for subsequent multimodal fusion, ensuring that the image information can be effectively represented in a form suitable for combination with other modal data.

[0069] S23. Design an innovative feature mapping scheme aimed at effectively docking the high-dimensional image vectors extracted from ViT with the text vectors optimized by Qwen. This requires developing a bimodal-compatible representation space that can not only retain the unique information of each modality but also promote the mutual understanding and fusion of the two;

[0070] The steps of the mapping scheme are as follows:

[0071] 1. Input the high-dimensional image vectors obtained from ViT and the text vectors optimized by Qwen into the feature mapping module, and map the feature vectors of the two modalities into a bimodal-compatible representation space; in this space, the image and text features are aligned and corresponding to each other, laying a foundation for subsequent fusion, and introducing semantic similarity calculation and context correlation analysis;

[0072] 2. In the mapped representation space, special attention is paid to achieving deep interweaving of image and text features at the semantic level, introducing advanced semantic analysis techniques, including semantic similarity calculation and context correlation analysis: First, semantic similarity calculation identifies the similar and related parts between the image and text features by calculating the semantic similarity between them; then, context correlation analysis fuses the context features of the image and text to ensure that the two can mutually confirm and complement each other at the semantic level;

[0073] 3. Stack the features fused at the semantic level to generate fused features;

[0074] 4. Use methods such as weighted summation, concatenation, or non-linear transformation to convert the fusion result of the image and text features into a unified fused feature representation, and finally output the multimodal fusion features after feature mapping and fusion processing.

[0075] These features not only retain the unique information of images and texts respectively, but also achieve the deep fusion and complementarity of the two modalities of information, forming a rich and accurate multi-modal context understanding. Based on the above fusion mechanism, a unified multi-modal representation framework is constructed. This framework encapsulates and standardizes the processed multi-modal information, forming a highly condensed and generalized representation form. This representation can not only promote the efficient interaction between modalities, but also provide a powerful basic platform for downstream tasks such as cross-modal retrieval, multi-modal generation, and context understanding, enhancing the generalization ability and task adaptability of the model.

[0076] Through the above specific processing process, the multi-modal feature mapping fusion mechanism can effectively fuse image and text features deeply, providing strong support and guarantee for multi-modal tasks.

[0077] S3. Carry out the refined training operation of the algorithm model. This process involves using a small amount of carefully selected high-quality power transmission business data sets to implement deep customized optimization of the pre-trained multi-modal large model. The specific strategies include using Lightweight Adapter (Lora Adaptive) and quantized Lora (Q-Lora) technologies. These two methods can maintain the stability of the original pre-trained model parameters while efficiently adjusting the model weights by introducing trainable low-rank matrices, realizing the rapid absorption and internalization of business domain knowledge.

[0078] Coarse tuning was carried out before, and now fine tuning is carried out. On the premise of not affecting the basic generality of the model, further perform full fine-tuning, which is a more refined parameter adjustment strategy. On the basis of keeping the overall architecture of the model unchanged, it gradually optimizes all parameters to ensure that the model can not only deeply understand the complex characteristics and industry specifications of the power transmission business, but also accurately respond to the subtle changes and specific needs in the business. This strategy not only improves the prediction accuracy and decision-making efficiency of the model in specific business scenarios, but also effectively avoids the risk of overfitting, ensuring the generalization performance and long-term application value of the model.

[0079] Specifically, the S3 specifically includes:

[0080] S31. Apply Lora Adaptive and Q-Lora for lightweight personalized adaptation.

[0081] First, select a representative small-scale high-quality data set from the massive power transmission business data through a screening algorithm to ensure that the data set covers the core scenarios and potential challenges of the business, providing accurate learning samples for model tuning. Preferably, the screening algorithm is a pre-trained visual model, ResNet50 or VGG16.

[0082] Introduce the Lightweight Adapter (Lora Adaptive) technology, add a low-rank matrix to the feature extraction layer of the model to achieve efficient parameter expansion without significantly increasing the computational burden. This step aims to inject specific business logic into the model without affecting its original generalization ability, accelerating the model's learning and adaptation to business knowledge.

[0083] To further optimize resource consumption and accelerate the training process, adopt the quantized version of Lora (Q-Lora) technology, add Lora at the decoder position of the multi-modal large model, and compress the model size by reducing the magnitude and bit width of the model parameters (such as changing from double to int8 type), while maintaining the model performance. This strategy improves the efficiency of training and inference while maintaining the tuning effect, especially suitable for resource-constrained environments.

[0084] During the application of Lora series technologies, pay special attention to the stability of the parameters of the original pre-trained model to ensure that the generality and stability of the model are not damaged while introducing personalized adjustments.

[0085] S32. Full-scale fine-tuning to achieve in-depth business customization:

[0086] During the full-scale fine-tuning process, adopt strategies such as regularization, early stopping, and data augmentation to effectively monitor and control the model complexity, avoid overfitting, and ensure that the model can perform well in specific business scenarios and maintain good generalization ability to unseen data.

[0087] Design a comprehensive and detailed full-scale fine-tuning plan, which includes determining key parameters such as the range of fine-tuned layers, learning rate strategy, batch size, etc., to ensure that the fine-tuning process is both meticulous and efficient and controllable. Perform the full-scale fine-tuning operation, and make detailed adjustments to all parameters while maintaining the stability of the model's macro architecture; and through iterative learning, gradually optimize each parameter to better match the specific needs of the power transmission business, and deeply explore and learn the complex patterns and specific rules in the business data.

[0088] S4. Perform model inference on the hidden danger identification algorithm obtained in the above steps to detect the hidden danger images of the power transmission channel.

[0089] Experimental data comparison or results.

[0090]

[0091] In this patent, by optimizing the algorithm model, the performance metrics of the model have been significantly improved. This improvement is not only reflected in the numerical increase, but also means that the model has stronger accuracy and robustness in practical applications. The improvement in precision indicates that the model is more accurate in identifying and classifying data. This model can provide correct results more reliably, reducing the risks and losses caused by misjudgments. The increase in recall means that the model can better capture all relevant data, reducing the situation of missed detections. Improving both precision and recall indicates that the model has achieved good results in balancing accuracy and comprehensiveness.

[0092] In practical applications, it is often necessary to trade off between precision and recall, and this model can find a better balance point between the two to meet the needs of different scenarios. This improvement also reflects the great potential of the optimized algorithm. By continuously optimizing the algorithm model, the performance of the model can be continuously improved, laying a solid foundation for future research and applications.

[0093] Embodiment 2

[0094] This embodiment provides a device for implementing an algorithm for the operation and maintenance of transmission lines based on a multi-modal pre-trained large model. The device includes:

[0095] At least one processor; and

[0096] A memory that stores instructions, which when executed by the at least one processor, cause the at least one processor to execute the algorithm for the operation and maintenance of transmission lines based on the multi-modal pre-trained large model as described above.

[0097] In this embodiment, the device includes but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, etc.

[0098] Embodiment 3

[0099] This embodiment also provides a computer-readable storage medium that stores executable instructions, which when executed cause the machine to execute the algorithm for the operation and maintenance of transmission lines based on the multi-modal pre-trained large model as described above.

[0100] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program codes for implementing the functions of any one of the above embodiments are stored, and the computer or processor of the system or device is made to read and execute the instructions stored in the readable storage medium.

[0101] In this case, the program code read from the readable medium itself can implement the functions of any one of the above-described embodiments. Therefore, the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0102] Examples of the readable storage medium include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or a cloud via a communication network.

[0103] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0104] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions specified in Figure 1One process or multiple processes and / or boxes Figure 1 Steps of the functions specified in one box or multiple boxes.

[0107] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A transmission line operation and maintenance algorithm based on a multimodal pre-trained large model, characterized in that: The method comprises: S1. Obtain a hidden danger image in a power transmission channel scenario, and attach text description information to the hidden danger in the image to form an image-text sample format; S2. Establish a multimodal pre-training large model, including a large-scale language model, a ViT image processing model and a feature mapping module; the large-scale language model and the ViT image processing model process the text data and the image data respectively, and then the feature mapping module maps and features the processed text vector and image vector; S3. Perform fine-tuning training on the established multimodal pre-trained large model, including lightweight personalized adaptation using Lora Adaptive and Q-Lora and designing fine-tuning parameters; S4. Train and verify the multimodal pre-trained large model, and then use it to detect hidden danger images of transmission channels.

2. The transmission line operation and maintenance algorithm based on multimodal pre-trained large model according to claim 1 is characterized in that: The S2 specifically includes: S21. Application and optimization of the large-scale language model Qwen: Based on the Encoder-Decoder architecture, the Encoder adopts the encoder structure of the large-scale language model Qwen, and the Decoder adopts the decoder of the large-scale language model Qwen; The input of the large-scale language model Qwen is preprocessed text data, and the preprocessing includes word segmentation, tokenization, and conversion into a format recognized by the model; The input text data first enters the encoder, which processes the input word vectors or character sequences one by one, captures the vocabulary relationship, grammatical structure and semantic information in the text through the encoder's feature extraction layer, and generates a high-dimensional encoding feature representation; then, the context information is fused with the current input text data; finally, it is input into the decoder for decoding to obtain the optimized text vector representation; S22, innovative integration of ViT image processing model and image feature extraction, where the encoder of ViT is stacked by multiple Transformer blocks, each of which includes multi-head self-attention and feed-forward neural networks, and residual connections and layer normalization are applied between each Transformer block; The input of the ViT image processing model is image data. The image is first segmented into multiple patches, and then the self-attention mechanism is applied to process the patches to capture the global dependencies and subtle features in the image. Then, the feature extraction layer of the encoder is used to perform in-depth analysis on the input image to obtain a high-dimensional image vector representation. S23, input the high-dimensional image vector obtained by ViT and the text vector after Qwen optimization into the feature mapping module, and map the feature vectors of the two modalities into a bimodal compatible representation space; in this space, the image and text features are aligned and corresponded to each other, and semantic similarity calculation and context association analysis are introduced: First, semantic similarity calculation calculates the semantic similarity between image and text features to identify the semantically similar and related parts of the two. Then, contextual association analysis fuses the contextual features of the image and text. Superimpose the features fused at the semantic level to generate fused features; The weighted summation method is adopted to convert the fusion results of image and text features into a unified fusion feature representation, and the final output is a multimodal fusion feature after feature mapping and fusion processing.

3. The transmission line operation and maintenance algorithm based on multimodal pre-trained large model according to claim 2 is characterized in that: The S3 specifically includes: S31. Use Lora Adaptive and Q-Lora for lightweight personalized adaptation: Select a small-scale, high-quality, representative dataset from the transmission business data; The Lora Adaptive technology is introduced to add low-rank matrices to each feature extraction layer of the multimodal pre-trained large model. At the same time, the quantized version of the Lora technology is used to add Lora to the decoder position of the multimodal large model to compress the model size by reducing the magnitude and bit width of the model parameters. S32, full-scale fine-tuning to achieve deep business customization: Design fine-tuning parameters and adopt regularization, early stopping, and data augmentation strategies during the full fine-tuning process to effectively monitor and control model complexity; the fine-tuning parameters include but are not limited to the range of layers, learning rate strategy, and batch size.

4. The transmission line operation and maintenance algorithm based on multimodal pre-trained large model according to claim 1 is characterized in that: The S1 specifically includes: S11, collecting and acquiring hidden danger images in the transmission channel scenario; S12. Manually attach text description information to the acquired image hidden dangers.

5. The transmission line operation and maintenance algorithm based on multimodal pre-trained large model according to claim 3 is characterized in that: The reduction of the magnitude and bit width of the model parameters is to convert double to int8 type.

6. A transmission line operation and maintenance device based on a multimodal pre-trained large model, characterized in that: The device comprises: processor; a memory having stored thereon a computer program executable on the processor; Wherein, when the computer program is executed by the processor, the steps of the transmission line operation and maintenance algorithm based on the multimodal pre-trained large model as described in any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Video abstract generation method based on large model fine tuning

    CN117412147A

  • Data integration method based on knowledge graph

    CN118333059A

  • Image search method based on multi-modal algorithm

    CN119226549A

  • Power transmission line operation and maintenance method and device based on multi-modal feature fusion and computer readable storage medium

    CN119339197A

Cited By

  • Power distribution network operation risk prompting method and system

    CN121073226A