Model modularization optimization method under multi-task learning
Through the model component optimization method under multi-task learning, the tensor decomposition and self-attention mechanism are used to disassemble the cross-modal large model into independent components, solving the limitations of traditional models in multi-task learning, achieving efficient processing of multi-modal data in the power system and intelligent decision-making support, and improving the operating efficiency and reliability of the power system.
Patent Information
- Application Number
- CN202510247692.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional single-task models are difficult to meet the needs of multiple complex power tasks. Cross-modal large models face disasters in multi-task learning, poor generalization performance, large computing power requirements, and insufficient application flexibility in different scenarios, and lack effective methods to collaborate with multi-tasks and optimize model structure.
The model component optimization method under multi-task learning is adopted to construct a cross-modal large model through tensor decomposition, and the weight is adjusted using self-attention mechanism and gradient calculation, and it is disassembled into multiple independent model components. Combined with Action-Driven component driving technology, the model is quickly customized and applied in different scenarios.
It improves the speed and accuracy of multimodal data processing, enhances the comprehensive analysis capabilities and decision-making support of the power system, optimizes resource allocation, and improves the operating efficiency and reliability of the power system.
Smart Images

Figure CN120297101A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and specifically to a method for optimizing model componentization under multi-task learning. Background Art
[0002] In the power industry, with the growth of data volume and the diversification of business requirements, a vast amount of power system production business data has emerged continuously, covering rich and diverse modal forms such as text, pictures, and videos. Traditional data analysis models have exposed many limitations when dealing with these complex multi-modal data, making it difficult to efficiently mine the internal relationships and values between data, and unable to meet the urgent needs of the power industry for accurate decision-making and efficient operation and maintenance.
[0003] Traditional single-task models are difficult to meet the requirements of simultaneously processing multiple complex power tasks. Cross-modal large models face problems such as network parameter dimensionality disasters, poor generalization performance, high computing power requirements, and insufficient flexibility in application in different scenarios during the multi-task learning process. In tasks such as power transmission and distribution equipment identification and substation equipment abnormal condition identification, multiple modal data need to be processed and there are correlations between tasks, but the existing technology lacks effective methods to coordinate these tasks and optimize the model structure to adapt to the deployment in different scenarios. Therefore, a method for optimizing model componentization under multi-task learning is designed to solve the above problems. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for optimizing model componentization under multi-task learning to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following solutions. The method for optimizing model componentization under multi-task learning includes the following steps:
[0006] A1. Widely collect multi-modal data from various data sources of the power system, covering image data, video data, and text data of power transmission and distribution equipment. Subsequently, perform preprocessing on different modal data. After the preprocessing is completed, store the preprocessed various modal data according to predetermined rules.
[0007] A2. Construct a cross-modal large model according to the requirements of power multi-task learning and the characteristics of multi-modal data. The core part of this model is to perform tensor decomposition on the model network parameters, express them as a combination of multiple low-rank tensors, and at the same time set the initial structure and sharing mode of the core tensor in the hidden space of the model, and then set the corresponding initial parameters for the model.
[0008] A3. Start model training using the preprocessed power multi-modal data. In each round of training, first calculate the gradient magnitude and direction of the power task with respect to the model parameters. Then, for the power transmission and distribution equipment defect detection task, obtain accurate gradient information by deriving and calculating the partial derivatives of the loss function with respect to the model parameters, so as to clearly reflect the learning dynamics and optimization direction of the model on this task. Then, according to the calculated gradient magnitude, determine the task weight adjustment coefficient through the weight adjustment function. Subsequently, use the self-attention mechanism to calculate the weights of the input power data of different modalities, and update the model parameters based on the gradient.
[0009] A4. When the model has undergone multiple rounds of training and its power performance reaches the expected set goal, disassemble the trained cross-modal large model according to functions and tasks into multiple independent model components, and perform a fine disassembly of the cross-modal large model according to the functional and task requirements. Subsequently, use the Action-Driven component driving technology to define a series of operations to drive the operation and collaboration of the model components. By formulating different Action sequences, flexibly call and combine the model components according to specific business requirements to achieve the rapid customization and application of the multi-task learning model in different scenarios.
[0010] In a further embodiment, the method of tensor decomposition is as follows:
[0011] A201. For different modalities of power multi-modal data, perform feature extraction and statistical data to determine the main feature directions and data distribution laws of each modality data in the high-dimensional space. According to the modal feature analysis results, construct a tensor structure suitable for power multi-modal data.
[0012] A202. Analyze the connections and data dependencies between multiple tasks in the power cross-modal large model, and then determine which tasks can share some tensor components through the correlation analysis of the task input and output data and the task execution process.
[0013] A203. Based on the task association recognition results, decompose the tensor data and optimize it to obtain a combined data of multiple low-rank tensors.
[0014] In a further embodiment, the preprocessed different modality data includes image data, video data, and text data;
[0015] The image data includes images of the overall appearance of the equipment, close-ups of key parts, and internal structure analysis. Its preprocessing is to adjust the contrast of the image to highlight image details, remove noise, purify the image, and perform intelligent cropping to highlight the key area, so that the image meets the model input requirements;
[0016] The video data includes the monitoring videos of the entire process of device operation and under different working conditions. Its preprocessing is to first extract key frames from the video data, and then perform the same preprocessing process on the key frames as that of the image data to obtain video data that meets the input requirements of the model.
[0017] The text data includes detailed device operation logs, maintenance records, operation specifications, and device usage guides. Its preprocessing is to perform comprehensive lexical and syntactic analysis on the text data, remove stop words, and use the method of text vectorization to convert the text into a vector form that can be processed by the model.
[0018] In a further embodiment, the method and its calculation formula for calculating the gradient magnitude of the power task with respect to the model parameters are as follows:
[0019] Input the preprocessed power multi-modal data according to the input layer structure of the model, then perform feature extraction on the input multi-modal data to obtain the feature vectors of each modal data. Then, fuse the features of different modalities. The fused feature vectors pass through the fully connected layer, and finally obtain the output of the model for each power task. Then, calculate the loss function for the output data.
[0020] Select the corresponding loss function according to the nature of the power task. Let the true device fault type label be U true , then the cross-entropy loss function K can be expressed as:
[0021]
[0022] Among them, M is the number of samples, and are the true label and predicted label of the l-th sample respectively;
[0023] For the regression task, the mean squared error loss function may be adopted;
[0024]
[0025] Among them, W is the mean squared error loss function, and are the true device performance index value and predicted value of the l-th sample respectively;
[0026] Through the above calculation of the loss function, obtain the gradient magnitude information of the model parameters.
[0027] In a further embodiment, the method and its calculation formula for calculating the gradient direction of the model parameters based on the loss function are as follows:
[0028] Perform backpropagation calculation based on the loss function. First, calculate the gradient of the loss function with respect to the model output Calculated according to the derivative formula of the selected loss function, which is the cross-entropy loss function;
[0029]
[0030] Then calculate the gradient of the model output with respect to the feature fusion vector of the previous layer, using the derivative calculation of the fully connected layer, etc., for the output of the fully connected layer;
[0031]
[0032] Among them, G fj is the feature fusion vector, H is the weight matrix, and T is the bias vector;
[0033] Subsequently, calculate the gradient data of the input of each layer with respect to the output of the previous layer in turn until the input data layer is calculated, and finally obtain the loss function with respect to each trainable parameter;
[0034]
[0035] Among them, ∈ is the gradient magnitude value of the trainable parameter. From this, the gradient magnitude and direction of each power task with respect to the model parameters can be obtained, where the magnitude of the gradient is the absolute value, and the direction is represented by its positive or negative sign.
[0036] In a further embodiment, the weight calculation of the input power data of different modalities is specifically as follows:
[0037] Let the image data of the multi-modal data be X T and the video data be X S and the text data be X W , first, convert the data of different modalities into feature representations through the corresponding embedding layers;
[0038] The image embedding is:
[0039] E T = CNN(X T )
[0040] Among them, E T is the representation of the image embedding, and features are extracted through the convolutional neural network;
[0041] The video embedding is:
[0042] E S = RNN(X S )
[0043] Among them, E S is the representation of the video embedding, and features are extracted through the recurrent neural network;
[0044] The text embedding is:
[0045] E W = WE(X W ) + PE(X W )
[0046] where WE is to convert the text into a word vector representation, and PE is to add position information;
[0047] Combine the embedding representations of different modalities to form a unified input matrix E, E = [E T , E S , E W , and then map E to the T, S, W matrices through a linear transformation;
[0048] T = Q T E
[0049] S = Q S E
[0050] T = Q W E
[0051] Z = S(Q T , Q S , Q W )
[0052] where Q T , Q S and Q S are the corresponding learned weight matrices, and Z is the data weight, so that the weights of different modalities can be determined.
[0053] In a further embodiment, the model components include but are not limited to a data preprocessing component, a feature extraction component, and a task-specific processing component.
[0054] In a further embodiment, the model components are encapsulated to form an independent reusable unit, and a model component database is established to classify, store, and manage the model components, and record the functions, input / output requirements, and dependency relationship information of the model components.
[0055] In a further embodiment, when the Action-Driven component receives a power service task request, it first parses the task to extract the key information of the task, which includes but is not limited to the task objective, the data modalities involved, and the processing accuracy, and then determines the corresponding Action sequence for completing the task according to the parsing result and the predefined Action classification and association rules.
[0056] In a further embodiment, after the parameters of the model component are quantized, the high-precision parameters in the model are converted into a low-precision data type. At the same time, in order to compensate for the accuracy loss caused by quantization, fine-tuning training is performed after quantization. A small amount of power data samples are used to retrain the quantized model, and the quantized parameters are adjusted to make the model restore its performance as much as possible while maintaining light weight. Subsequently, by analyzing the importance of the parameters in each layer of the model, the connections or neurons that have less impact on the model performance are removed, realizing the light weight of the multi-task model componentization.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0058] The power cross-modal large model of the present invention can effectively integrate multi-modal data such as text, pictures, and videos in the production business of the power system, break through the data modality barriers of traditional models, and realize in-depth correlation analysis of data through the self-attention mechanism, greatly improving the speed and accuracy of data processing, and providing timely and reliable support for the daily operation and decision-making of the power industry; with comprehensive analysis capabilities and automatic report generation functions, this model can provide comprehensive and in-depth data analysis reports and intelligent decision-making suggestions for the managers of the power industry; by mining and analyzing multi-modal data, it helps managers discover potential problems, predict equipment failures, and power supply and demand trends in a timely manner, thereby optimizing resource allocation and improving the overall operation efficiency and reliability of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a flowchart of the method for optimizing the model componentization of the present invention;
[0060] Figure 2 It is a flowchart of the method for tensor decomposition of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] Embodiment 1
[0063] Please refer to FIGS. 1-2. The method for optimizing the model componentization under multi-task learning provided in this embodiment includes the following steps:
[0064] A1. Widely collect multimodal data from various data sources in the power system. These data cover image data, video data, and text data of power transmission and distribution equipment. The image data includes images of the overall appearance of the equipment, close-ups of key parts, and dissections of internal structures, which can intuitively reflect the physical state of the equipment. The video data records the entire process of equipment operation and monitoring videos under different working conditions, helping to dynamically observe the equipment operation. The text data includes detailed equipment operation logs, maintenance records, operation specifications, and equipment usage guides, providing a basis for the historical information and operation guidelines of the equipment.
[0065] Subsequently, preprocess different modal data. For image data, adjust the contrast of the image to highlight image details, making the features of the equipment clearer and more distinguishable, which helps the model to more accurately identify the key parts and potential defects of the equipment, and improve the model's ability to analyze image information. Remove noise and purify the image to reduce the impact of interference factors on the model's judgment, avoid misjudgment caused by noise interference, and enhance the stability and reliability of the model. Use intelligent cropping to highlight the key area, making the image fit the model input requirements, improving the model processing efficiency, reducing unnecessary data processing volume, and saving computing resources and time costs. For video data, first extract key frames, which can represent the main information of the video, and then perform the same preprocessing process on the key frames as the image data to obtain video data that meets the model input requirements, not only effectively reducing the complexity of video data processing, but also ensuring that the model can obtain key dynamic information from the video. For text data, conduct comprehensive lexical and syntactic analysis, remove stop words, and use text vectorization to convert the text into a vector form that can be processed by the model, enabling the model to understand and process text information, laying a foundation for the model to fuse multimodal data for comprehensive analysis, and helping the model to mine important information such as equipment operation status and maintenance conditions from the text. After the preprocessing is completed, store the preprocessed various modal data according to predetermined rules for convenient subsequent calling and use, improving the convenience of data management, avoiding data chaos, and ensuring that the required data can be obtained quickly and accurately during model training and application.
[0066] A2: Build a cross-modal large model
[0067] According to the requirements of power multi-task learning and the characteristics of multi-modal data, a cross-modal large model is constructed. In the core part of the model, tensor decomposition is performed on the model network parameters. Specifically, first, for different modalities of power multi-modal data, feature extraction and statistical data are carried out to determine the main feature directions and data distribution laws of each modality data in the high-dimensional space. According to the results of modality feature analysis, a tensor structure suitable for power multi-modal data is constructed. This targeted tensor structure construction can make full use of the characteristics of each modality data, effectively integrate different modality information, provide strong support for the model to achieve cross-modal learning, and improve the model's comprehensive processing ability for complex power data. Then, the connections and data dependencies between multiple tasks in the power cross-modal large model are analyzed. Through the correlation analysis of the input and output data of the tasks and the task execution process, it is determined which tasks can share some tensor components. Based on the task association recognition results, the tensor data is decomposed and optimized, and the model network parameters are expressed as a combination form of multiple low-rank tensors, reducing the complexity of the model parameters, reducing the amount of calculation, improving the efficiency of model training and inference. At the same time, by sharing tensor components, the information interaction and collaboration between different tasks are enhanced, and the overall performance of the model in multi-task processing is improved. At the same time, the initial structure and sharing mode of the core tensor in the hidden space are set for the model, and then the corresponding initial parameters are set for the model, providing a good start for the subsequent stable training and efficient learning of the model, ensuring that the model can be optimized and iterated on a reasonable parameter basis;
[0068] A3: Model Training
[0069] Model training starts with the preprocessed power multi-modal data. In each round of training, first, the gradient magnitude and direction of the power task with respect to the model parameters are calculated. For the power transmission and distribution equipment defect detection task, by deriving and calculating the partial derivative of the loss function with respect to the model parameters, accurate gradient information is obtained, which clearly reflects the learning dynamics and optimization direction of the model on this task, enabling the model to accurately adjust the parameters according to the gradient information, accelerating convergence, and improving the learning efficiency and accuracy of the model for the defect detection task. Then, according to the calculated gradient amount, the task weight adjustment coefficient is determined through the weight adjustment function, which can dynamically adjust the importance of different tasks in model training, enabling the model to reasonably allocate learning resources in the multi-task learning process, avoiding some tasks from overly dominating the training process, ensuring that each task can be fully learned, and improving the generalization ability of the model in multi-task scenarios. Subsequently, the self-attention mechanism is used to calculate the weights of the input power data of different modalities, and the model parameters are updated based on the gradient. The self-attention mechanism can enable the model to automatically focus on the most important parts of different modality data for the current task, enhancing the model's ability to capture key information. Combining the gradient to update the parameters helps the model to be quickly optimized under complex multi-modal data, improving the adaptability and performance of the model;
[0070] A4: Model Disassembly and Application
[0071] After the model undergoes multiple rounds of training and its power performance reaches the expected set goal, the trained cross-modal large model is disassembled according to functions and tasks, divided into multiple independent model components, and finely disassembled according to the functional and task requirements. This disassembly method makes the model have higher flexibility and scalability. Each model component focuses on a specific function or task, facilitating subsequent maintenance, upgrading, and optimization of individual components. Subsequently, using the Action-Driven component-driven technology, a series of operations are defined to drive the operation and collaboration of model components. By formulating different Action sequences, model components can be flexibly invoked and combined according to specific business requirements, realizing the rapid customization and application of the multi-task learning model in different scenarios. This greatly improves the application efficiency of the model, can quickly build appropriate model solutions for different power business scenarios, reduce the development cycle and cost, meet the diverse and personalized business needs of the power industry, and enhance the operation and management capabilities of the power system in various complex scenarios.
[0072] Example 2
[0073] Refer to Figure 1-2 , and further improvements are made on the basis of Example 1:
[0074] The method and its calculation formula for calculating the gradient magnitude of the power task with respect to the model parameters are as follows:
[0075] The preprocessed power multi-modal data is input according to the input layer structure of the model, and then feature extraction is performed on the input multi-modal data to obtain the feature vectors of each modal data. Then, the features of different modalities are fused. The fused feature vectors pass through the fully connected layer, and finally the output of the model for each power task is obtained. Then, the loss function is calculated for the output data;
[0076] Select the corresponding loss function according to the nature of the power task. Let the true device fault type label be U true , then the cross-entropy loss function K can be expressed as:
[0077]
[0078] where M is the number of samples, and are the true label and predicted label of the l-th sample respectively;
[0079] For regression tasks, the mean squared error loss function may be adopted;
[0080]
[0081] Among them, \(W\) is the mean squared error loss function, and are the true device performance index value and the predicted value of the \(l\)-th sample respectively. Through the above calculation of the loss function, the gradient magnitude information of the model parameters is obtained, providing a quantitative basis for the model to update parameters based on the gradient, ensuring that the model can adjust parameters along the direction most conducive to reducing the loss during training, thereby improving the training effect and convergence speed of the model;
[0082] Through the above calculation of the loss function, the gradient magnitude information of the model parameters is obtained.
[0083] The method and its calculation formula for calculating the gradient direction of the model parameters based on the loss function are as follows:
[0084] Perform backpropagation calculation based on the loss function. First, calculate the gradient of the loss function with respect to the model output Calculate according to the derivative formula of the selected loss function, for the cross-entropy loss function;
[0085]
[0086] Then calculate the gradient of the model output with respect to the feature fusion vector of the previous layer, using the derivative calculation of the fully connected layer, etc., for the output of the fully connected layer;
[0087]
[0088] Among them, \(G\) fj is the feature fusion vector, \(H\) is the weight matrix, and \(T\) is the bias vector;
[0089] Subsequently, calculate the gradient data of the input of each layer with respect to the output of the previous layer in turn until the input data layer is calculated, and finally obtain the loss function with respect to each trainable parameter;
[0090]
[0091] Among them, \(\epsilon\) is the gradient magnitude value of the trainable parameter. Thus, the gradient magnitude and direction of each power task with respect to the model parameters can be obtained. Among them, the magnitude of the gradient is the absolute value, and the direction is represented by its positive or negative sign. This precise way of calculating the gradient direction enables the model to optimize along the most effective path in the parameter space, avoiding the model falling into local optimal solutions during training, improving the global search ability of the model, ensuring that the model can converge to the optimal parameters faster and more stably, and enhancing the generalization performance of the model.
[0092] The weight calculation for the input power data of different modalities is specifically as follows:
[0093] Let the image data of the multimodal data be X T and the video data be X S and the text data be X W , first, convert the data of different modalities into feature representations through the corresponding embedding layers;
[0094] The image embedding is:
[0095] E T = CNN(X T )
[0096] where E T is the representation of the image embedding, and the features are extracted through a convolutional neural network. The convolutional neural network has unique advantages in processing image data and can effectively extract the local features of the image, providing guarantee for the subsequent accurate analysis of the device information in the image;
[0097] The video embedding is:
[0098] E S = PNN(X S )
[0099] where E S is the representation of the video embedding, and the features are extracted through a recurrent neural network. The recurrent neural network is suitable for processing video data with sequence characteristics and can capture the time series information in the video, enabling the model to better understand the dynamic process of device operation;
[0100] The text embedding is:
[0101] E W = WE(X W ) + PE(X W )
[0102] where WE converts the text into a word vector representation, and PE adds position information. By the above embedding method, the problem of the order information of words in the text is solved. In this way, the semantic and structural information in the text data is fully mined, providing rich text features for the model to comprehensively analyze multimodal data;
[0103] Combine the embedding representations of different modalities to form a unified input matrix E, E = [E T , E S , E W , then, map E to the T, S, W matrices through a linear transformation;
[0104] T = Q T E
[0105] S = Q S E
[0106] T = QW E
[0107] Z = S(Q T ,Q S ,Q W )
[0108] where Q T 、Q S and Q S are weight matrices corresponding to learning, and Z is the data weight, so that the weights of different modalities can be confirmed.
[0109] This way of weight calculation can adaptively adjust the importance of different modality data in the model, enabling the model to reasonably allocate the attention to each modality data according to the task requirements, improving the accuracy and effectiveness of the model's multi-modal data fusion processing, and enhancing the model's performance in comprehensively analyzing multi-source information of power equipment.
[0110] Example 3
[0111] Referring to Figure 1-2 , further improvements are made on the basis of Example 1:
[0112] The model components include but are not limited to data preprocessing components, feature extraction components, and task-specific processing components. The data preprocessing components are responsible for preprocessing the original multi-modal data as described in Example 1 to ensure that the data quality and format meet the processing requirements of subsequent components, providing a reliable data basis for the efficient operation of subsequent components, avoiding the decline of model performance caused by data quality problems. Precise feature extraction can greatly improve the model's ability to understand and process data, enhancing the model's discriminative ability. The task-specific processing components then perform targeted processing on the extracted features according to specific power tasks to achieve the task objectives, enabling the model to efficiently complete various specific power tasks and improving the applicability of the model in different power business scenarios.
[0113] The model components are encapsulated to make them an independent and reusable unit, and a model component database is established to classify, store, and manage the model components, recording the functions, input / output requirements, and dependency relationship information of the model components. By encapsulating the model components and establishing a database, the reusability of the model components is improved, reducing repetitive development work and development costs. At the same time, it is convenient for developers to quickly search for and call the required components, improving development efficiency. It is also convenient for unified management and maintenance of the model components, ensuring the stability and reliability of the model components.
[0114] The Action-Driven component-driven receives a power business task request. First, it parses the task to extract the key information of the task. The key information includes but is not limited to the task objective, the involved data modalities, and the processing accuracy. Then, according to the parsing results and the predefined Action classification and association rules, it determines the Action sequence corresponding to completing this task. For example, if the task objective is to detect a specific fault type of power transmission and distribution equipment, the involved data modalities may be image data and text data, and the processing accuracy requirement may be to achieve a certain accuracy rate and recall rate. By accurately parsing the task key information, it can provide an accurate basis for determining the appropriate Action sequence later, ensuring that the called model components and the executed operations highly match the task requirements.
[0115] Based on the Action classification and association rules, it determines the Action sequence corresponding to completing this task. Action classification can be divided according to the nature and operation type of the task, such as data processing Action, model training Action, result evaluation Action, etc. The association rules define the sequence and dependency relationships between different Actions. By reasonably determining the Action sequence, it can efficiently call and combine model components to complete power business tasks. This rule-driven component call method improves the flexibility and efficiency of model component combination, can quickly respond to different power business needs, and provides strong support for the intelligent operation of the power system.
[0116] After quantifying the parameters of the model component, it converts the high-precision parameters in the model into low-precision data types, such as converting 32-bit floating-point numbers to 16-bit floating-point numbers or 8-bit integers. This can significantly reduce the storage requirements and computational amount of the model, improve the running efficiency of the model in resource-constrained environments. For example, the model can run more smoothly on edge devices or mobile devices, expanding the application scenarios of the model. At the same time, to compensate for the accuracy loss caused by quantization, fine-tuning training is performed after quantization. Using a small amount of power data samples to retrain the quantized model and adjust the quantized parameters, so that the model can restore performance as much as possible while maintaining lightweight, ensuring that the model still has high accuracy and reliability under the premise of lightweight.
[0117] Subsequently, by analyzing the importance of the parameters in each layer of the model, it removes the connections or neurons that have less impact on the model performance, further realizing the lightweight of the multi-task model component. The importance of the parameters can be evaluated, for example, by calculating the gradient magnitude of the parameters or using pruning algorithms. In this way, without affecting the main performance of the model, it reduces the complexity of the model, improves the running speed and resource utilization efficiency of the model, enabling the model to run more efficiently under limited hardware resources and providing better support for application scenarios with high real-time requirements in the power system.
[0118] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for optimizing model componentization under multi-task learning, characterized in that, It includes the following steps: A1. Widely collect multimodal data from various data sources of the power system, covering image data, video data, and text data of power transmission and distribution equipment. Subsequently, preprocess different modalities of data. After the preprocessing is completed, store the preprocessed various modalities of data according to predetermined rules; A2. According to the power multi-task learning requirements and the characteristics of multimodal data, construct a cross-modal large model. The core part of this model is to perform tensor decomposition on the model network parameters, represent them in the form of a combination of multiple low-rank tensors, and at the same time set the initial structure and sharing mode of the core tensor in the hidden space of this model, and then set the corresponding initial parameters for the model; A3. Use the preprocessed power multimodal data to start model training. In each round of training, first calculate the magnitude and direction of the gradient of the power task with respect to the model parameters. Then, for the power transmission and distribution equipment defect detection task, obtain accurate gradient information by deriving and calculating the partial derivative of the loss function with respect to the model parameters, so as to clearly reflect the learning dynamics and optimization direction of the model on this task. Then, according to the calculated gradient amount, determine the task weight adjustment coefficient through the weight adjustment function. Subsequently, use the self-attention mechanism to calculate the weights of the input power data of different modalities, and update the model parameters based on the gradient; A4. When the model has undergone multiple rounds of training and the power performance reaches the expected set goal, disassemble the trained cross-modal large model according to functions and tasks, divide it into multiple independent model components, and perform a fine disassembly of the cross-modal large model according to the functional and task requirements. Subsequently, use the Action-Driven component driving technology to define a series of operations to drive the operation and collaboration of the model components. By formulating different Action sequences, flexibly call and combine the model components according to specific business requirements to achieve the rapid customization and application of the multi-task learning model in different scenarios.
2. The model component optimization method under multi-task learning according to claim 1, wherein The method of the tensor decomposition is specifically as follows: A201. For different modalities of power multimodal data, perform feature extraction and statistical data, determine the main feature directions and data distribution laws of each modality of data in the high-dimensional space, and construct a tensor structure suitable for power multimodal data according to the modal feature analysis results; A202. Analyze the connections and data dependencies between multiple tasks in the power cross-modal large model, and then determine which tasks can share some tensor components through the correlation analysis of the input and output data of the tasks and the task execution process; A203. Based on the task association recognition results, decompose the tensor data and optimize it to obtain a combined data of multiple low-rank tensors.
3. The method for optimizing model componentization under multi-task learning according to claim 1, characterized in that, The preprocessed different modalities of data include image data, video data, and text data; The image data includes images of the overall appearance of the equipment, close-ups of key parts, and internal structure dissections. Its preprocessing is to adjust the contrast of the image to highlight image details, remove noise, purify the image, and perform intelligent cropping to highlight the key area, so that the image meets the model input requirements; The video data includes monitoring videos of the entire process of equipment operation and under different working conditions. Its preprocessing is to first extract key frames from the video data, and then perform the same preprocessing process on the key frames as that of the image data to obtain video data that meets the requirements of the model input. The text data includes detailed equipment operation logs, maintenance records, operation specifications, and equipment user guides. Its preprocessing is to perform comprehensive lexical and syntactic analysis on the text data, remove stop words, and use the method of text vectorization to convert the text into a vector form that can be processed by the model.
4. The method for optimizing model componentization under multi-task learning according to claim 1, wherein, The method and its calculation formula for calculating the gradient magnitude of the computing power task with respect to the model parameters are as follows: Input the preprocessed multi-modal power data according to the input layer structure of the model, then perform feature extraction on the input multi-modal data to obtain the feature vectors of each modal data, then fuse the features of different modalities. The fused feature vectors pass through the fully connected layer, and finally the output of the model for each power task is obtained, and then the loss function is calculated for the output data. Select the corresponding loss function according to the nature of the power task. Let the true device fault type label be U true , then the cross-entropy loss function K can be expressed as: where M is the number of samples, and are the true label and the predicted label of the l-th sample, respectively; For the regression task, the mean squared error loss function may be adopted. where W is the mean squared error loss function, and are the true device performance metric value and the predicted value of the l-th sample, respectively; Through the above calculation of the loss function, the gradient magnitude information of the model parameters is obtained.
5. The method for optimizing model componentization under multi-task learning according to claim 4, characterized in that The method and its calculation formula for calculating the gradient direction of the model parameters based on the loss function are as follows: Perform backpropagation calculation based on the loss function. First, calculate the gradient of the loss function with respect to the model output Calculate according to the derivative formula of the selected loss function, which for the cross-entropy loss function; Then calculate the gradient of the model output with respect to the feature fusion vector of the previous layer, and use the derivative calculation of the fully connected layer, etc., for the output of the fully connected layer. Among them, G fj is the feature fusion vector, H is the weight matrix, and T is the bias vector; Subsequently, calculate the gradient data of the input of each layer with respect to the output of the previous layer in turn until the input data layer is calculated, and finally obtain the loss function with respect to each trainable parameter. Among them, ∈ is the gradient magnitude value of the trainable parameters, from which the gradient magnitude and direction of each power task with respect to the model parameters can be obtained. The magnitude of the gradient is the absolute value, and the direction is represented by its positive or negative sign.
6. The method for optimizing model componentization under multi-task learning according to claim 1, characterized in that The specific weight calculation of the input power data of different modalities is as follows: Let the image data of the multimodal data be X T and the video data be X S and the text data be X w , first, convert the data of different modalities into feature representations through corresponding embedding layers; The image embedding is: E T = CNN(X T ) Among them, E T is the performance of image embedding, and features are extracted through a convolutional neural network; The video embedding is: E S = RNN(X S ) Among them, E S is the performance of video embedding, and features are extracted through a recurrent neural network; The text embedding is: E W = WE(X W ) + PE(X W ) Wherein, WE is to convert the text into a word vector representation, and PE is to add position information. Combine the embedding representations of different modalities to form a unified input matrix E, where E = [E T , E S , E W . Then, map E to the T, S, W matrices through a linear transformation; T = Q T E S = Q S E T = Q W E Z = S(Q T ,Q S ,Q W ) Among them, Q T , Q S and Q S are weight matrices corresponding to learning, and Z is the data weight, so that the weights of different modalities can be confirmed.
7. The method for optimizing model componentization under multi-task learning according to claim 1, characterized in that The model components include but are not limited to data preprocessing components, feature extraction components, and task-specific processing components.
8. The method for optimizing model componentization under multi-task learning according to claim 1, wherein Package the model components to make them an independent reusable unit, and establish a model component database to classify, store, and manage the model components, and record the functions, input / output requirements, and dependency information of the model components.
9. The method for optimizing model componentization under multi-task learning according to claim 1, characterized in that The Action-Driven component drives to receive a power business task request. First, it parses the task to extract the key information of the task. The key information includes but is not limited to the task objective, the involved data modalities, and the processing accuracy. Then, according to the parsing result and the predefined Action classification and association rules, it determines the Action sequence corresponding to completing the task.
10. The method for optimizing model componentization under multi-task learning according to claim 1, wherein After the parameters of the model components are quantized, the high-precision parameters in the model are converted into low-precision data types. At the same time, in order to compensate for the accuracy loss caused by quantization, fine-tuning training is performed after quantization. Use a small amount of power data samples to retrain the quantized model, adjust the quantized parameters, so that the model can restore performance as much as possible while maintaining lightweight. Subsequently, by analyzing the importance of the parameters of each layer in the model, remove the connections or neurons that have less impact on the model performance to achieve the lightweight of the multi-task model components.