Power distribution network equipment health state analysis method and device based on multi-modal large model

By constructing a multimodal large model and combining hierarchical learning rate adjustment and mixed precision training, the compatibility and accuracy issues of multimodal data processing in the health status assessment of distribution network equipment were solved, enabling a comprehensive and accurate analysis of equipment health status and improving the accuracy and efficiency of the assessment.

CN120995363APending Publication Date: 2025-11-21STATE GRID HEBEI ELECTRIC POWER RES INST +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510684334.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing methods for assessing the health status of distribution network equipment are mainly limited to a single data source, making it difficult to comprehensively and accurately reflect the actual operating status of the equipment. In particular, they suffer from compatibility and accuracy issues when processing multimodal data, and cannot meet the high-precision assessment requirements of smart grids.

Method used

A multimodal large model is adopted. By acquiring text, image and audio data of power distribution network equipment, a model fine-tuning dataset is constructed. The parameters of the preset multimodal large model are fine-tuned by combining hierarchical learning rate adjustment and mixed precision training. A multimodal evaluation model is constructed, and the model parameters are optimized by using the whale algorithm to achieve a comprehensive and accurate analysis of the health status of the equipment.

Benefits of technology

It enables a comprehensive and accurate analysis of the health status of power distribution network equipment, improves the accuracy and efficiency of assessment, and can promptly detect equipment aging or faults, ensuring the safe and healthy operation of equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995363A_ABST
    Figure CN120995363A_ABST
Patent Text Reader

Abstract

The invention provides a power distribution network equipment health state analysis method and device based on a multi-modal large model, and relates to the technical field of power grids. According to the method, multi-modal data such as text data, image data and audio data of power distribution network equipment are comprehensively analyzed, a model fine tuning data set is constructed in combination with the state type of the equipment, and parameter fine tuning is performed on a trained multi-modal large model. In the fine adjustment process, a hierarchical learning rate adjustment and mixed precision training mode is adopted, model parameters are finely adjusted in combination with an evaluation function, and prediction of the multi-modal evaluation model obtained through adjustment is more comprehensive and accurate. Compared with an evaluation model adopting a single data source, the multi-modal evaluation model can comprehensively and accurately analyze the health state of the power distribution network, various data sources do not need to be detected respectively, and the accuracy and efficiency of health state analysis of the power distribution network equipment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power grid, and particularly relates to a power distribution network equipment health state analysis method and device based on a multi-modal large model. BACKGROUND

[0002] Under the background of the current intelligent power grid technology changing with each passing day, the health state evaluation of power distribution network equipment has become a core link to ensure the safe and stable operation of the power system. As the key infrastructure for power transmission and distribution, the accurate evaluation of the running state of power distribution network equipment plays an irreplaceable role in preventing faults, improving power supply reliability and optimizing resource allocation. However, the traditional equipment health state evaluation method is mostly limited to the analysis of a single data source, and this limitation makes the evaluation results often unable to comprehensively and truly reflect the actual running state of the equipment.

[0003] With the continuous development of Internet of Things, big data and artificial intelligence technology, power distribution network equipment generates a large amount of multi-modal data during operation, which covers various types of information such as electrical parameters, temperature, vibration, image and video. However, the existing evaluation methods face serious challenges in integrating and processing these multi-modal data. On the one hand, the differences between different modal data are significant, and the compatibility and consistency between models are difficult to guarantee, which greatly increases the difficulty of data fusion. On the other hand, due to the limitations of model design, the existing methods often have low evaluation accuracy when dealing with complex and nonlinear multi-modal data, which is difficult to meet the demand of intelligent power grid for high-precision evaluation of equipment health state. SUMMARY

[0004] The present application provides a power distribution network equipment health state analysis method and device based on a multi-modal large model, which can comprehensively and accurately analyze the health state of the power distribution network, and improve the accuracy and efficiency of the power distribution network equipment health state analysis.

[0005] In a first aspect, the present application provides a power distribution network equipment health state analysis method based on a multi-modal large model, which comprises: acquiring multi-modal data of the power distribution network equipment, the multi-modal data including text data, image data and audio data; based on the multi-modal data and each state type of the equipment, constructing a model fine-tuning data set, the state type including equipment health, slight damage and serious damage; based on the model fine-tuning data set, an evaluation function pre-constructed, using a hierarchical learning rate adjustment and mixed precision training method, fine-tuning the parameters of a pre-set multi-modal large model to obtain a multi-modal evaluation model; based on the multi-modal evaluation model, performing health state analysis on the power distribution network equipment.

[0006] In a possible implementation, the model fine-tuning data set is constructed based on the multi-modal data and the state types of the device, including: labeling the multi-modal data based on the state types of the device to obtain an initial data set of the state types; and performing data augmentation on the multi-modal data in the initial data set to obtain the model fine-tuning data set, the data augmentation including one or more of the following: partial replacement, random exchange, random insertion, and random deletion.

[0007] In a possible implementation, the parameters of the preset multi-modal large model are fine-tuned based on the model fine-tuning data set, a pre-constructed evaluation function, a hierarchical learning rate adjustment, and mixed precision training to obtain a multi-modal evaluation model, including: step one, performing hierarchical learning rate adjustment on the preset multi-modal large model to obtain a first model, the hierarchical learning rate adjustment including dividing levels, setting learning rates, and configuring optimizers; step two, performing mixed precision training based on the model fine-tuning data set and the first model to obtain a second model; step three, calculating the gradient of the second model and adjusting the parameters of the second model based on the gradient to obtain a fine-tuned model; step four, evaluating the fine-tuned model based on the evaluation function to obtain an evaluation result; step five, if the evaluation result meets the performance requirement, the fine-tuned model is determined as the multi-modal evaluation model; and step six, if the evaluation result does not meet the performance requirement, steps one to six are repeatedly executed until the evaluation result meets the performance requirement.

[0008] In a possible implementation, before the parameters of the preset multi-modal large model are fine-tuned based on the model fine-tuning data set, a pre-constructed evaluation function, a hierarchical learning rate adjustment, and mixed precision training to obtain a multi-modal evaluation model, the method further includes: constructing an accuracy function of each modality data; constructing a weight function of each modality data based on the accuracy function of each modality data; constructing a comprehensive accuracy function based on the accuracy function and the weight function of each modality data; constructing an energy consumption function based on an energy consumption parameter, the energy consumption parameter including a training convergence time, a model parameter size, and a floating point operation amount; and constructing the evaluation function based on the comprehensive accuracy function and the energy consumption function.

[0009] In one possible implementation, a health status analysis of distribution network equipment is performed based on a multimodal assessment model, including: acquiring real-time multimodal data of the distribution network equipment; inputting the real-time multimodal data into the multimodal assessment model to obtain the output results and the assessment process parameters of the multimodal assessment model; the assessment process parameters include recognition accuracy, recognition time, and computational cost; calculating the function value of the objective value function based on the assessment process parameters of the multimodal assessment model; adjusting the model parameters of the multimodal assessment model using the whale algorithm with the objective value function as the goal, to obtain the adjusted multimodal assessment model; obtaining the optimal output result based on the adjusted multimodal assessment model and the real-time multimodal data; and determining the health status type of the distribution network equipment based on the optimal output result.

[0010] Secondly, embodiments of the present invention provide a device for analyzing the health status of distribution network equipment based on a multimodal large model. The device includes a communication module and a processing module. The communication module is used to acquire multimodal data of the distribution network equipment, including text data, image data, and audio data. The processing module is used to construct a model fine-tuning dataset based on the multimodal data and various state types of the equipment, including healthy equipment, slightly damaged equipment, and severely damaged equipment. Based on the model fine-tuning dataset and a pre-constructed evaluation function, the parameters of a pre-defined multimodal large model are fine-tuned using hierarchical learning rate adjustment and mixed precision training to obtain a multimodal evaluation model. Based on the multimodal evaluation model, the health status of the distribution network equipment is analyzed.

[0011] In one possible implementation, the processing module is specifically used to label the multimodal data based on the various state types of the device to obtain an initial dataset for each state type; and to perform data augmentation on the multimodal data in the initial dataset to obtain a model fine-tuning dataset. The data augmentation includes one or more of the following: partial replacement, random swapping, random insertion, and random deletion.

[0012] Thirdly, embodiments of the present invention provide an electronic device including a memory and a processor. The memory stores a computer program, and the processor is configured to call and run the computer program stored in the memory to perform the steps of the method as described in the first aspect and any possible implementation thereof.

[0013] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method as described in the first aspect and any possible implementation thereof.

[0014] The application provides a power distribution network equipment health state analysis method and device based on a multi-modal large model, and the application comprehensively analyzes multi-modal data such as text data, image data and audio data of power distribution network equipment, constructs a model fine-tuning data set in combination with the state type of the equipment, and fine-tunes parameters of a trained multi-modal large model. In the fine-tuning process, a hierarchical learning rate adjustment and mixed precision training mode are adopted, the model parameters are fine-tuned in combination with an evaluation function, the prediction of the adjusted multi-modal evaluation model is more comprehensive and accurate. Compared with an evaluation model using a single data source, the multi-modal evaluation model in the application can comprehensively and accurately analyze the health state of the power distribution network state, does not need to detect various data sources respectively, and improves the accuracy and efficiency of power distribution network equipment health state analysis. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0016] Figure 1 is a flowchart of a power distribution network equipment health state analysis method based on a multi-modal large model provided by the embodiments of the present application;

[0017] Figure 2 is a flowchart of model full amount fine-tuning provided by the embodiments of the present application;

[0018] Figure 3 is a structural schematic diagram of a power distribution network equipment health state analysis device based on a multi-modal large model provided by the embodiments of the present application;

[0019] Figure 4 is a structural schematic diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0020] In the following description, specific details such as specific system structures, techniques, etc. are presented in order to thoroughly understand the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details that hinder the description of the present application.

[0021] In the description of the application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this paper is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can mean: A alone, A and B exist at the same time, and B alone. In addition, "at least one" "multiple" means two or more. "First", "second" and the like do not limit the quantity and execution order, and "first", "second" and the like do not necessarily mean different.

[0022] In the embodiments of the present application, the words such as "exemplary" or "for example" are used to mean an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. On the contrary, the use of "exemplary" or "for example" is intended to present the relevant concept in a specific manner, so as to facilitate understanding.

[0023] In addition, the terms "include" and "have" mentioned in the description of the present application and any modification thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or modules is not limited to the listed steps or modules, but can optionally include other steps or modules not listed, or can optionally include other steps or modules inherent to the process, method, product or device.

[0024] In order to make the purpose, technical scheme and advantages of the present application clearer, specific embodiments will be described below with reference to the accompanying drawings of the present application.

[0025] As described in the background, the current health state evaluation of power distribution network equipment has the technical problems of complex data mode, difficult processing, and low processing precision.

[0026] To solve the above technical problems, the application provides a power distribution network equipment health state analysis method based on a multi-modal large model. Multi-modal data such as operation parameters, vibration, temperature and images of the power distribution network equipment are obtained by means of current, voltage, vibration, temperature and other multi-type sensor monitoring and unmanned aerial vehicle inspection photographing; a pre-trained multi-modal large model is selected based on an open source large model on the network; a fine-tuning data set is made based on the data obtained in the early stage, the model is fine-tuned in full amount based on the layer-wise learning rate decay (Layer-Wise Learning Rate Decay) and mixed precision training (Mixed Precision Training) method, an evaluation function is constructed to evaluate the performance of the fine-tuned model, and a double-layer model is used to strengthen the learning of the large model. The trained large model is applied to evaluate the health state of the power distribution network equipment, to timely find the health degradation state such as equipment aging and failure, and to develop a reasonable maintenance strategy to intervene in advance, so as to ensure the safe and healthy operation of the power distribution network equipment.

[0027] As shown in Figure 1 The embodiment of the application provides a power distribution network equipment health state analysis method based on a multi-modal large model. The method comprises steps S101-S104.

[0028] S101, obtaining multi-modal data of the power distribution network equipment.

[0029] In the embodiment of the application, the multi-modal data includes text data, image data and audio data.

[0030] For example, the embodiment of the application can obtain multi-modal data such as text, image and audio of the power distribution network equipment based on sensors, unmanned aerial vehicles and artificial inspection, and pre-process the data through data cleaning.

[0031] S102, constructing a model fine-tuning data set based on the multi-modal data and the state types of the equipment.

[0032] In the embodiment of the application, the state types include equipment health, slight damage and serious damage.

[0033] As a possible implementation manner, step S102 can be implemented as steps S1021-S1022.

[0034] S1021, labeling the multi-modal data based on the state types of the equipment to obtain an initial data set of each state type.

[0035] S1022, data augmentation is performed on the multi-modal data in the initial data set to obtain the model fine-tuning data set.

[0036] In some embodiments, the data augmentation includes one or more of partial replacement, random exchange, random insertion, and random deletion.

[0037] Exemplarily, the embodiment of the application can use ChatGPT-4.0 as a pre-trained multi-modal large model, perform data augmentation on small-scale but high-quality artificially annotated power distribution network equipment damage state data through methods such as partial replacement, random exchange, random insertion, and random deletion of text, images, and audio, and make an enhanced power grid equipment damage state training set. The system needs to make a classification judgment among (“device health”, “minor damage”, “severe damage”).

[0038] The classification category set is: C = {C1, C2, C3}, wherein:

[0039] Device health (C1): The device is in a normal working state, with no abnormal performance, and no damage signs are detected in the monitoring data, such as current, voltage, temperature, and other parameters are within the normal working interval, and no repair is needed.

[0040] Minor damage (C2): The device function is still available, but there are abnormalities in some places, such as current, temperature deviating from the normal working interval at some time or mechanical component surface light wear, etc., which need to be included in the repair plan.

[0041] Severe damage (C3): The device is severely damaged, which may affect its normal operation, such as current, temperature deviating from the normal working interval for a long time, mechanical component protection layer separation, etc., which needs to be repaired immediately.

[0042] For the text modality, there are N D test samples and corresponding true labels Each prediction result is For the image and audio modalities, the same applies, there are N and

[0043] S103, based on the model fine-tuning dataset, the pre-constructed evaluation function, using hierarchical learning rate adjustment and mixed precision training, fine-tuning the parameters of the pre-set multi-modal large model to obtain a multi-modal evaluation model.

[0044] In some embodiments, the embodiment of the application can download an open source large model, select a multi-modal large model that has been pre-trained in the energy and power field, and access the power distribution network equipment health state online monitoring system

[0045] As a possible implementation, as shown in Figure 2 , step S103 can be implemented as steps one to six.

[0046] Step one, adjust the learning rate of the preset multi-modal large model by layer, and obtain a first model.

[0047] In some embodiments, the hierarchical learning rate adjustment includes dividing layers, setting learning rates, and configuring optimizers.

[0048] For example, to optimize the adaptability of the pre-trained model to the power distribution network equipment health assessment task, the present application sets different learning rate ratios for different layers of the model to adjust the model parameters.

[0049] 1. Layer division

[0050] The pre-trained model M is divided into multiple layers L1, L2,..Ln. Where n is the number of layers, the bottom layer is responsible for extracting general features of equipment health status, the middle layer is for feature combination, and the top layer is adapted to the demand of equipment health status evaluation.

[0051] 2. Set learning rate

[0052] Set different learning rate ratios a for each layer Li i , where a i ∈(0,1], to control the update amplitude of each layer of parameters. The base learning rate is set to η, then the learning rate of each layer is η i The calculation formula is as follows:

[0053] η i = a i η;

[0054] 3. Optimizer configuration

[0055] In the optimizer, configure the corresponding parameter group θ i for each layer Li, and apply the corresponding learning rate η i . The parameter update formula of the optimizer is:

[0056]

[0057] Where, is the parameter of layer Li at the t-th iteration; is the loss function of the gradient of the parameter θ i .

[0058] Step two, based on the model fine-tuning dataset and the first model, perform mixed precision training to obtain a second model.

[0059] For example, considering the large number of model parameters and long calculation time, the method of mixed precision training is adopted, using half-precision (FP16) and single-precision (FP32) floating-point calculation to optimize the calculation resources in the training process.

[0060] 1. Automatic Mixed Precision Management

[0061] During training, mixed precision computation is adopted, converting part of the computationally intensive tasks to half-precision (FP16) execution, while keeping the learning rate and other key parameter updates in single-precision (FP32). The representation precision of the weight matrix W is set as:

[0062] W FP16 = convert_to_FP16(W FP32 );

[0063] where convert_to_FP16 represents the conversion of single-precision weights to half-precision representation.

[0064] 2. Loss Scaling

[0065] To prevent the gradient value from being too small (underflow) or too large (overflow) in half-precision calculation, a dynamic loss scaling factor s is introduced, s ∈ (0.5, 10], and after each backpropagation step, the loss scaling factor is dynamically adjusted based on the maximum value of the gradient to adjust the scaling of the loss function :

[0066]

[0067] During backpropagation, the gradient calculation is:

[0068]

[0069] Then, the gradient is recovered by reverse scaling:

[0070]

[0071] 3. Gradient Calculation and Update

[0072] During the forward and backward propagation stages, mainly use FP16 for calculation to improve the calculation speed and reduce the memory occupation:

[0073] Forward Pass: y = f(WFP16, x)

[0074] Backward Pass:

[0075] During the weight update stage, FP32 is used for parameter update to ensure the accuracy and stability of the parameters:

[0076]

[0077] Step three, calculate the gradient of the second model, and based on the gradient of the second model, adjust the parameters of the second model to obtain the fine-tuned model.

[0078] Step four, evaluating the fine-tuned model based on the evaluation function to obtain an evaluation result;

[0079] Step five, if the evaluation result meets the performance requirement, determining the fine-tuned model as the multi-modal evaluation model;

[0080] Step six, if the evaluation result does not meet the performance requirement, repeating steps one to six until the evaluation result meets the performance requirement.

[0081] S104, performing health state analysis on the power distribution network equipment based on the multi-modal evaluation model.

[0082] As a possible implementation, step S104 can be implemented as steps S1041-S1046.

[0083] S1041, obtaining real-time multi-modal data of the power distribution network equipment.

[0084] S1042, inputting the real-time multi-modal data into the multi-modal evaluation model to obtain an output result and evaluation process parameters of the multi-modal evaluation model.

[0085] In some embodiments, the evaluation process parameters include recognition accuracy, recognition time and computational complexity.

[0086] S1043, calculating a function value of the target value function based on the evaluation process parameters of the multi-modal evaluation model.

[0087] In some embodiments, to continuously improve the multi-modal large model equipment health state evaluation capability, a double-layer model reinforcement learning strategy is established in combination with the Actor-Critic algorithm to optimize the large model. First, a positive and negative reward feedback model is constructed based on the Critic algorithm, then the large model parameters are optimized according to the feedback of the Critic, and finally the weight parameters in the feedback model are optimized using the whale algorithm according to the focus requirements of reality on recognition accuracy, time and computational complexity, to ensure that the equipment health state evaluation capability of the multi-modal large model meets the needs.

[0088] In some embodiments, the upper layer positive and negative reward feedback model can be established based on the Critic algorithm and the model parameters are updated according to the revenue. For example, the target value function is determined by the following formula:

[0089] R = ω1·R accuracy + ω2·R time + ω3·R computational_load ;

[0090] Wherein, R is the function value of the target value function, i.e. the final revenue after performing the action in the current state, R accuracyTo identify the accuracy reward, R time To identify the time reward, R computational_load To calculate the amount of reward, ω1 is the weight of the identification accuracy reward, ω2 is the weight of the identification time reward, and ω3 is the weight of the calculation amount reward.

[0091] For example, the identification accuracy reward can be determined by the following formula.

[0092]

[0093] In the formula, α1 and β1 are preset coefficients; ACC represents the identification accuracy; and ACCthreshold represents the identification accuracy threshold, which needs to be comprehensively formulated based on the computing resources and the scene tasks to avoid excessive single optimization or sacrifice of a certain aspect of the model in order to improve the total score. When the accuracy exceeds the threshold, the reward increases slowly with the increase of the accuracy. When the accuracy is lower than the threshold, the punishment increases in the form of square, and the punishment degree gradually increases with the decrease of the accuracy.

[0094] For example, the identification time reward can be determined by the following formula.

[0095]

[0096] In the formula, α2 and β2 are preset coefficients, T is the identification time, and Tthreshold is the identification time threshold, which needs to be comprehensively formulated based on the computing resources and the scene tasks to avoid excessive single optimization or sacrifice of a certain aspect of the model in order to improve the total score. When the identification time is lower than the threshold, the punishment increases slowly with the increase of the identification time. When the identification time exceeds the threshold, the punishment increases in the form of exponential, and the punishment degree rapidly increases with the increase of the identification time.

[0097] For example, the calculation amount reward can be determined by the following formula.

[0098]

[0099] In the formula, C is the calculation amount, and Cthreshold is the calculation amount threshold, which needs to be comprehensively formulated based on the computing resources and the scene tasks to avoid excessive single optimization or sacrifice of a certain aspect of the model in order to improve the total score. When the calculation amount is less than the threshold, the reward increases slowly with the decrease of the calculation amount. When the calculation amount exceeds the threshold, the punishment increases exponentially with the increase of the calculation amount, ensuring that the calculation amount is in a reasonable range to improve the overall performance of the system.

[0100] The feedback model is optimized according to the Critic to build the model parameter θ. The update formula is:

[0101]

[0102] wherein, θ represents a model parameter, and a represents a learning rate ratio, The policy gradient is represented. The delta represents a time difference (TD) error, which measures the deviation between the value prediction of the current state and the actual observed reward and the value of the next state, and the formula is:

[0103] delta = R + gamma * V π (s') - V π (s).

[0104] wherein, gamma represents a discount factor, which is initially 1 and is dynamically set based on the current and next state value, and is used to balance the influence of the current reward and the future reward, gamma * V π (s') represents the value estimation of the next state s', and gamma * V π (s) represents the value estimation of the current state s.

[0105] S1044, taking the optimal target value function as the goal, adjusting the model parameters of the multi-modal evaluation model by using the whale optimization algorithm to obtain an adjusted multi-modal evaluation model.

[0106] For example, the embodiment of the application can establish a lower whale optimization model to optimize the parameters in the upper model, dynamically adjust the weight of the upper feedback function by using the whale algorithm, simulate the hunting method of humpback whales, assume that the position of the optimal whale in a certain population is the position of the target prey, and consider that the probabilities of selecting the hunting method of searching for a surrounding or spiral bubble net to update the position of the whale in the optimization process are each 50%. That is, there is a random number r1 between [0, 1].

[0107] When r1 is greater than or equal to 0.5, the hunting method of the spiral bubble net is used for position updating, and the position updating is:

[0108] Z(t+1,j) = Le bλ cos(2 * pi) + Z * (t,j)

[0109] In the formula, Z(t,j) is the position of the whale in the j-dimensional variable of the tth generation population; Z*(t,j) is the position corresponding to the j variable of the whale that obtains the global optimal solution in the tth generation population; L is the distance between the whale and the prey, L = |Z * (t,j) - z(t,j)|; b is a constant that controls the shape of the logarithmic spiral; and lambda is a random number in [-1, 1].

[0110] When p is less than 0.5, the hunting method of searching for a surrounding is used for position updating, and whether it is a prey searching stage or a surrounding stage is determined according to the modulus value of the coefficient vector A.

[0111] When |A| is less than 1, the prey surrounding stage is entered, and the position updating is:

[0112]

[0113] wherein, L' is the step length of the surrounding; A and C are the coefficient vectors respectively; r2 is a random number between [0, 1]; a is a convergence factor, which decreases linearly from 2 to 0 with the increase of the iteration number; Tmax is the maximum iteration number.

[0114] When |A|≥1, the prey search stage is entered, and the position is updated as:

[0115]

[0116] wherein, Z rand (t) is the position of the j-th variable randomly selected from the current population; L'' is the search step length.

[0117] Exemplarily, the embodiment of the present application can apply the reinforcement learning strategy to the multi-modal large model, continuously optimize and adjust the power distribution network equipment health state evaluation, evaluate the equipment health state according to the accepted text, audio and image data, quickly determine the current state of the equipment, and formulate the corresponding maintenance plan.

[0118] S1045, based on the adjusted multi-modal evaluation model and real-time multi-modal data, an optimal output result is obtained.

[0119] S1046, based on the optimal output result, the health state type of the power distribution network equipment is determined.

[0120] The present application provides a power distribution network equipment health state analysis method based on a multi-modal large model, which comprehensively analyzes multi-modal data such as text data, image data and audio data of the power distribution network equipment, and constructs a model fine-tuning data set combined with the state type of the equipment to fine-tune the parameters of the trained multi-modal large model. In the fine-tuning process, hierarchical learning rate adjustment and mixed precision training are adopted, and the model parameters are fine-tuned combined with the evaluation function. The prediction of the adjusted multi-modal evaluation model is more comprehensive and accurate. Compared with the evaluation model using a single data source, the multi-modal evaluation model in the present application can comprehensively and accurately analyze the health state of the power distribution network state, without the need for various data sources to be detected respectively, thereby improving the accuracy and efficiency of the power distribution network equipment health state analysis.

[0121] The application creates a model fine-tuning data set according to multi-modal data such as text, images and audio, uses a hierarchical learning rate adjustment and mixed precision training method to fine-tune the model in full, constructs an evaluation function to evaluate the fine-tuned model results, ensures the performance improvement of the iterative model, and uses a double-layer model to strengthen the learning of the large model, constructs a multi-modal large model for power distribution network equipment health state evaluation, realizes online evaluation of equipment health state and reasonable maintenance strategy formulation, maximizes resource utilization, minimizes fault impact and optimizes operation stability.

[0122] Optionally, the multi-modal large model-based power distribution network equipment health state analysis method provided by the embodiment of the application further comprises steps S201-S204 before step S103.

[0123] S201, construct an accuracy function of each modal data.

[0124] In some embodiments, the embodiment of the application can define the health state evaluation capability Eassess of the power distribution network equipment as the output-to-input ratio, set the minimum evaluation capability score according to the actual requirement for the accuracy of equipment health state recognition, and update the large model parameters immediately when the minimum score is not reached,

[0125] The equipment health state evaluation accuracy Q of text, image and audio represents the output, and the energy consumption ratio C represents the input, i.e., the energy cost required to provide the corresponding service, and the mathematical model is as follows:

[0126]

[0127] Wherein, D, I and A respectively represent the accuracy of the model in evaluating the health state of the equipment after inputting text, image and audio files; T represents the response time, which represents the model operation time. P represents the parameter size, which represents the model size and complexity. The more parameters, the more powerful the model, but the more computing resources and energy consumption; the floating point operation amount F directly affects the model running calculation amount; and λ represents a constant factor, which is initially 1 and is dynamically adjusted according to different requirements for the equipment evaluation state capability in different scenarios. The higher the requirement, the lower the constant factor, the lower the current health state evaluation capability score, and the greater the possibility of large model optimization parameters. The constant factor adjusts the health state evaluation capability value range.

[0128] In some embodiments, taking the text modal data as an example, the accuracy function of each modal data in the text modal data is determined by the following formula:

[0129]

[0130] Wherein, Q(D) is the accuracy of the text modal data, is the probability that the prediction value of the i-th test sample in the text modal data is the real value, N D is the number of test samples of the text modal data, is the prediction value of the i-th test sample in the text modal data, is the real value of the i-th test sample in the text modal data, and the image and audio modal accuracy can be listed in the same way.

[0131] S202, based on the accuracy function of each modal data, constructing the weight function of each modal data.

[0132] For example, considering the influence characteristics of text, audio and image accuracy, the weight is defined as:

[0133]

[0134] When predicting the final multi-modal fusion model, the model will receive text, video and audio inputs for each test sample (denoted as Xi): For this sample, the prediction probability distribution of three states of three modalities is obtained during training:

[0135]

[0136] Using the above defined weight, the weighted sum of the probabilities of the three modalities is obtained to obtain the fused prediction rate:

[0137]

[0138] Then, the class corresponding to the maximum probability in the fused probability distribution is taken as the final prediction class:

[0139]

[0140] The comprehensive accuracy is

[0141]

[0142] wherein,

[0143] S203, based on the accuracy function and the weight function of each modal data, constructing a comprehensive accuracy function.

[0144] For example, the accuracy function and the weight function of each modal data can be weighted and summed to determine the comprehensive accuracy function in the embodiments of the present application.

[0145] S204, constructing an energy consumption function based on an energy consumption parameter.

[0146] In some embodiments, the energy consumption parameter includes training convergence time, model parameter size and floating point operation amount.

[0147] The energy consumption function is determined by the following formula:

[0148]

[0149] wherein C(T, P, F) is the energy consumption function, represents the relative influence value of the training convergence time on the total energy consumption, represents the relative influence value of the model parameter size on the total energy consumption, represents the relative influence value of the floating point operation amount on the total energy consumption, T represents the efficiency coefficient of the training convergence time, P represents the efficiency coefficient of the model parameter size, F represents the efficiency coefficient of the floating point operation amount, which is set according to the requirements of different scenarios for evaluating the ability of large models in various aspects, and is 1 without special preference.

[0150] S205, constructing an evaluation function based on the comprehensive accuracy function and the energy consumption function.

[0151] In this way, the embodiment of the present application can evaluate the parameter adjustment process of the multi-modal model by constructing the evaluation function, thereby improving the accuracy of the adjusted multi-modal evaluation model.

[0152] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0153] The following is a device embodiment of the present application. For details not described in detail, reference can be made to the corresponding method embodiments described above.

[0154] Figure 3 A structure schematic diagram of a power distribution network equipment health state analysis device based on a multi-modal large model is shown. The analysis device 300 includes a communication module 301 and a processing module 302.

[0155] The communication module 301 is used to obtain multi-modal data of the power distribution network equipment, and the multi-modal data includes text data, image data and audio data.

[0156] The processing module 302 is used to construct a model fine-tuning data set based on the multi-modal data and each state type of the equipment, the state type including equipment health, slight damage and serious damage; fine-tune the parameters of a preset multi-modal large model in a hierarchical learning rate adjustment and mixed precision training manner based on the model fine-tuning data set, the pre-constructed evaluation function, to obtain a multi-modal evaluation model; and analyze the health state of the power distribution network equipment based on the multi-modal evaluation model.

[0157] Figure 4 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. As shown in Figure 4 the electronic device 400 includes a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. The processor 401 implements the steps in each method embodiment described above when executing the computer program 403, for example, steps S101-S104 shown in Figure 1 . Alternatively, the processor 401 implements the functions of each module / unit in each device embodiment described above when executing the computer program 403, for example, the functions of the communication module 301 and the processing module 302 shown in Figure 3 .

[0158] Illustratively, the computer program 403 can be divided into one or more modules / units, which are stored in the memory 402 and executed by the processor 401 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 403 in the electronic device 400. For example, the computer program 403 can be divided into Figure 3 the communication module 301 and the processing module 302 shown in .

[0159] The processor 401 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0160] The memory 402 can be an internal storage unit of the electronic device 400, such as a hard disk or a memory of the electronic device 400. The memory 402 can also be an external storage device of the electronic device 400, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like equipped on the electronic device 400. Further, the memory 402 can include both an internal storage unit and an external storage device of the electronic device 400. The memory 402 is used to store the computer program and other programs and data required by the terminal. The memory 402 can also be used to temporarily store data that has been output or will be output.

[0161] The above-described embodiments are merely intended for describing the technical solutions of the present application, but not to limit the present application; even though the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for analyzing the health status of distribution network equipment based on a multimodal large model, characterized in that, The method includes: Acquire multimodal data from power distribution network equipment, wherein the multimodal data includes text data, image data, and audio data; Based on the multimodal data and the various state types of the device, a model fine-tuning dataset is constructed, wherein the state types include device health, minor damage, and severe damage; Based on the model fine-tuning dataset, the pre-constructed evaluation function is fine-tuned using hierarchical learning rate adjustment and mixed precision training to obtain the parameters of the pre-set multimodal large model, thus obtaining the multimodal evaluation model. Based on the aforementioned multimodal assessment model, a health status analysis is performed on the distribution network equipment.

2. The method for analyzing the health status of distribution network equipment based on a multimodal large model according to claim 1, characterized in that, The process of constructing a model fine-tuning dataset based on the multimodal data and the various state types of the device includes: Based on the various state types of the device, the multimodal data is labeled to obtain the initial dataset for each state type; Data augmentation is performed on the multimodal data in the initial dataset to obtain a model fine-tuning dataset. The data augmentation includes one or more of the following: partial replacement, random swapping, random insertion, and random deletion.

3. The method for analyzing the health status of distribution network equipment based on a multimodal large model according to claim 1, characterized in that, The evaluation function, pre-constructed based on the model fine-tuning dataset, employs hierarchical learning rate adjustment and mixed precision training to fine-tune the parameters of the pre-defined multimodal large model, resulting in a multimodal evaluation model, including: Step 1: Perform hierarchical learning rate adjustment on the preset multimodal large model to obtain the first model. The hierarchical learning rate adjustment includes dividing into layers, setting the learning rate, and configuring the optimizer. Step 2: Based on the model fine-tuning dataset and the first model, perform mixed-precision training to obtain the second model; Step 3: Calculate the gradient of the second model, and adjust the parameters of the second model based on the gradient of the second model to obtain the fine-tuned model; Step four: Based on the evaluation function, evaluate the fine-tuned model to obtain the evaluation result; Step 5: If the evaluation results meet the performance requirements, then the fine-tuned model is determined as the multimodal evaluation model. Step six: If the evaluation result does not meet the performance requirements, repeat steps one through six until the evaluation result meets the performance requirements.

4. The method for analyzing the health status of distribution network equipment based on a multimodal large model according to claim 1, characterized in that, Before the pre-constructed evaluation function, based on the model fine-tuning dataset, is fine-tuned using hierarchical learning rate adjustment and mixed precision training to fine-tune the parameters of the pre-defined multimodal large model to obtain the multimodal evaluation model, the process further includes: Construct accuracy functions for each modality of data; Based on the accuracy function of each modality data, a weight function for each modality data is constructed; Based on the accuracy function and weighting function of each modality data, a comprehensive accuracy function is constructed. Based on energy consumption parameters, an energy consumption function is constructed, wherein the energy consumption parameters include training convergence time, model parameter size, and floating-point operation volume; The evaluation function is constructed based on the comprehensive accuracy function and the energy consumption function.

5. The method for analyzing the health status of distribution network equipment based on a multimodal large model according to claim 1, characterized in that, The accuracy function of the text modality data in each modality is determined by the following formula: Where Q(D) represents the accuracy of the text modal data. Let N be the probability that the predicted value of the i-th test sample in the text modal data is the true value. D This represents the number of test samples for the text modal data. Let be the predicted value of the i-th test sample in the text modal data. is the true value of the i-th test sample in the text modal data.

6. The method for analyzing the health status of distribution network equipment based on a multimodal large model according to claim 1, characterized in that, The energy consumption function is determined by the following formula: Where C(T,P,F) is the energy consumption function. This represents the relative impact of training convergence time on total energy consumption. This represents the relative impact of the model parameter size on the total energy consumption. A represents the relative impact of floating-point operations on total energy consumption. T A represents the efficiency coefficient for training convergence time. P The efficiency coefficient, A, represents the model parameter scale. F The efficiency coefficient representing the amount of floating-point operations.

7. The method for analyzing the health status of distribution network equipment based on a multimodal large model according to claim 1, characterized in that, The health status analysis of distribution network equipment based on the multimodal assessment model includes: Acquire real-time multimodal data of power distribution network equipment; The real-time multimodal data is input into the multimodal evaluation model to obtain the output results and the evaluation process parameters of the multimodal evaluation model; the evaluation process parameters include recognition accuracy, recognition time, and computational cost; Based on the evaluation process parameters of the multimodal evaluation model, the function value of the target value function is calculated; With the goal of optimizing the objective value function, the whale algorithm is used to adjust the model parameters of the multimodal evaluation model, resulting in the adjusted multimodal evaluation model. Based on the adjusted multimodal evaluation model and the real-time multimodal data, the optimal output result is obtained; Based on the optimal output result, the health status type of the distribution network equipment is determined.

8. The method for analyzing the health status of distribution network equipment based on a multimodal large model according to claim 1, characterized in that, The target value function is determined by the following formula: R=ω1·R accuracy +ω2·R time +ω3·R computational_load ; Where R is the value of the objective value function, R accuracy As a reward for recognition accuracy, R time To identify time rewards, R computational_load The reward is calculated based on computational effort, ω1 is the weight of the recognition accuracy reward, ω2 is the weight of the recognition time reward, and ω3 is the weight of the computational effort reward.

9. A device for analyzing the health status of distribution network equipment based on a multimodal large model, characterized in that, include: A communication module is used to acquire multimodal data from power distribution network equipment, wherein the multimodal data includes text data, image data, and audio data; The processing module is used to construct a model fine-tuning dataset based on the multimodal data and the various state types of the device, wherein the state types include device health, minor damage, and severe damage; Based on the model fine-tuning dataset, the pre-constructed evaluation function is fine-tuned using hierarchical learning rate adjustment and mixed precision training to obtain the parameters of the preset multimodal large model, thus obtaining the multimodal evaluation model; based on the multimodal evaluation model, the health status analysis of the distribution network equipment is performed.

10. The distribution network equipment health status analysis device based on a multimodal large model according to claim 9, characterized in that, The processing module is specifically used to label the multimodal data based on the various state types of the device to obtain an initial dataset for each state type; and to perform data augmentation on the multimodal data in the initial dataset to obtain a model fine-tuning dataset. The data augmentation includes one or more of the following: partial replacement, random swapping, random insertion, and random deletion.