Method and system for fine tuning and lightweight design of electric power visual large model

By constructing the power scene data set and fine-tuning and quantitatively optimizing the power vision model, the problem of insufficient data adaptability and poor flexibility of the power vision model in the power scene is solved, and efficient and low-cost intelligent power management is achieved.

CN120046683APending Publication Date: 2025-05-27NARI INFORMATION & COMM TECH +2
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411925250.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing power vision model optimization methods have problems such as insufficient data adaptability, poor flexibility, low quantization accuracy, and how to efficiently build power scenario data sets and realize intelligent power management with lightweight models.

Method used

The power scene data set is constructed based on the preprocessed power image data, and the power vision big model is fine-tuned, and the quantized post-training optimization method is used to compress it. Combined with knowledge distillation to optimize the model parameters, the optimal lightweight power vision big model is output.

Benefits of technology

The generalization ability of the power vision model in specific power scenarios is improved, the false alarm and missed alarm rates are reduced, the efficient compression and performance stability of the model are achieved, and the power industry's demand for high efficiency, low cost and reliability is met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046683A_ABST
    Figure CN120046683A_ABST
Patent Text Reader

Abstract

The invention discloses a fine tuning and lightweight design method and system for an electric power visual large model, and relates to the technical field of electric power visual large model optimizing.The method comprises the steps that an electric power scene data set is constructed based on preprocessed electric power image data; performing fine tuning on the electric power visual large model based on the electric power scene data set; and compressing the fine-tuned electric power vision large model based on a quantized post-training optimization method, optimizing parameters of the electric power vision large model through knowledge distillation, and outputting an optimal lightweight electric power vision large model. According to the method, fine tuning is performed on the electric power visual large model based on the electric power scene data set, the specific adaptive capacity of the electric power visual large model to the electric power scene is enhanced through a hierarchical fine tuning strategy, and network weight quantization is performed on the fine-tuned electric power visual large model through a quantized post-training optimization method. Compression of an electric power vision large model is achieved, model parameters are optimized through knowledge distillation, and model lightweight and efficient reasoning are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power vision large model optimization, and specifically to a method and system for fine-tuning and lightweight design of a power vision large model. Background Art

[0002] The power system plays a crucial role in modern society, and its operating efficiency and safety directly affect economic development and social stability. However, with the continuous expansion of the power grid scale and the continuous improvement of complexity, traditional power management methods can no longer meet the high-efficiency operation requirements of modern power grids. In key areas such as power equipment operation monitoring, fault troubleshooting, and intelligent decision-making, the existing technologies have exposed many deficiencies. In recent years, power image recognition technology based on computer vision has gradually emerged, realizing the intelligentization of power grid monitoring through the use of deep learning models. However, such technologies usually rely on small deep learning models, with limited model generalization ability and strong dependence on computing resources, making it difficult to effectively handle diverse and complex power scenarios. In addition, the classifiers and detectors of these models often need to be retrained when dealing with new tasks, resulting in a long development cycle and difficulty in quickly adapting to the dynamic changes of the power grid.

[0003] Although large-scale pre-trained vision models have attracted much attention due to their excellent feature extraction capabilities, when actually applied to power scenarios, they still face multiple challenges. First, the data characteristics of power scenarios have specificity and scarcity, and the data diversity is relatively low. Existing pre-trained models are mainly trained based on general scenario datasets and are difficult to accurately capture specific damage patterns or abnormal operation behaviors of power equipment. This lack of data adaptability leads to the performance of the model in power scenarios not meeting expectations. Second, pre-trained vision models usually have a huge parameter scale, often in the billions, and running and reasoning require high-end hardware support, which is too costly for many power enterprises and severely limits the feasibility of their promotion and application.

[0004] In addition, in terms of model compression and lightweight design, it is difficult for existing technologies to find a reasonable balance between performance and accuracy. Compression operations often lead to a significant decrease in model accuracy. Especially in tasks such as power equipment fault detection, the increase in misrecognition rate may bring serious safety hazards. At the same time, the complexity and diversity of power scenarios require the model to have the ability to flexibly adapt to different tasks. However, existing large-scale models often need to be fully retrained when switching tasks, with a long development cycle and poor flexibility. These deficiencies limit the practical application of large-scale pre-trained vision models in power scenarios and are difficult to meet the requirements of the power industry for high efficiency, low cost, and reliability. Summary of the Invention

[0005] In view of the above existing problems, the present invention is proposed.

[0006] Therefore, the technical problems solved by the present invention are: the existing optimization methods for power vision models have insufficient data adaptability, poor flexibility, low quantization accuracy, and the problems of how to efficiently construct a power scenario dataset and achieve intelligent power management with a lightweight model.

[0007] To solve the above technical problems, the present invention provides the following technical solutions: a method for fine-tuning and lightweight design of a power vision large model, including constructing a power scenario dataset based on the preprocessed power image data; fine-tuning the power vision large model based on the power scenario dataset; compressing the fine-tuned power vision large model by a post-training optimization method based on quantization, optimizing the parameters of the power vision large model through knowledge distillation, and outputting the optimal lightweight power vision large model.

[0008] As a preferred embodiment of the method for fine-tuning and lightweight design of the power vision large model of the present invention, wherein: the preprocessing of the power image data includes collecting power image data through a vision sensor at different times, locations, angles, lighting conditions, weather conditions, and seasons, including normal samples and abnormal samples of power equipment; the preprocessing includes cleaning the collected power image data, deleting blurred, duplicate, or irrelevant images, annotating the cleaned data, and further performing data augmentation on the annotated data.

[0009] As a preferred embodiment of the method for fine-tuning and lightweight design of the power vision large model of the present invention, wherein: the construction of the power scenario dataset includes dividing the power scenario dataset into a training set, a validation set, and a test set, the training set is used to train the power vision large model, the validation set is used to determine the optimal power vision large model, and the performance of the optimal power vision large model is tested through the test set.

[0010] As a preferred embodiment of the method for fine-tuning and lightweight design of the power vision large model of the present invention, wherein: the fine-tuning of the power vision large model includes fine-tuning the power vision large model based on the training set in the power scenario dataset, adopting a hierarchical fine-tuning strategy, initializing the feature extractor of the pre-trained vision large model, loading the vision large model dedicated to the power industry, fine-tuning or fully parameter fine-tuning by adjusting the freeze layer selection parameters of the pre-trained vision large model feature extractor, the feature extractor retains the low-level basic features of the pre-trained model, and thaws the high-level parameters layer by layer, while re-training the high-level classifier and detector; when the data features and class distributions of the task scenario are similar to those of the pre-trained model, partial parameter fine-tuning is performed on the feature extractor, and only the high-level is thawed for fine-tuning; when the task scenario contains new device types, brand-new categories, or different fault features, which are quite different from the pre-trained model, then full parameter fine-tuning is performed to thaw all network layers.

[0011] As a preferred solution of the method for fine-tuning and lightweight design of the power vision large model described in the present invention, it further includes: designing a detection head, where the designed detection head includes a detection head designed for the power scenario. The detection head is designed to select a two-stage detector or a single-stage detector according to the target detection task, and a total loss function is constructed through a classification loss function and a regression loss function; the cross-entropy loss is used to calculate the classification error, and the classification loss function is constructed, expressed as:

[0012]

[0013] where L cls is the classification loss function, y i is the true label of class i, and p i is the predicted probability of class i by the power vision large model; the smooth L1 loss is used to calculate the bounding box regression error, and the regression loss function is constructed, expressed as:

[0014]

[0015] where L reg is the regression loss function, and x is the error between the true bounding box position t and the predicted position q; the total loss function is calculated, expressed as:

[0016] L total = L cls + λL reg

[0017] where L total is the total loss function, and λ is the balancing factor.

[0018] As a preferred solution of the method for fine-tuning and lightweight design of the power vision large model described in the present invention, where: the compression includes network weight quantization of the fine-tuned power vision large model. Through the post-training optimization technology of quantization, the floating-point weights of the fine-tuned power vision model are converted into a low-precision integer format; the post-training optimization technology of quantization includes adaptive batch quantization with dynamic column combination. Adaptive batch quantization with dynamic column combination includes contribution degree evaluation, dynamic column grouping, and error feedback adjustment; contribution degree evaluation includes dividing each column vector of the weight matrix based on the locality-sensitive hashing technology and calculating the sensitivity index of each column weight in the model prediction accuracy; dynamic column grouping includes grouping columns according to the sensitivity index, and preferentially selecting column combinations with high contribution degree and close error distribution; error feedback adjustment includes, after each batch of quantization is completed, re-evaluating the contribution degree of the remaining columns according to the updated quantized column weights and dynamically adjusting the subsequent grouping strategy.

[0019] As a preferred solution of the method for fine-tuning and lightweight design of the power vision large model described in the present invention, wherein: optimizing the parameters of the power vision large model through knowledge distillation and outputting the optimal lightweight power vision large model includes constructing a teacher model and a student model, using the power vision large model after fine-tuning and compression as the teacher model, and passing the soft labels generated by the teacher model to the student model. By introducing the distillation temperature to smooth the prediction results, it is expressed as:

[0020]

[0021] wherein, y i is the prediction result of class i, z i is the original score of class i, K is the distillation temperature, zj is the original scores of all classes j; the prediction result after the teacher model is warmed up is the soft label, the prediction result after the student model is warmed up is the soft prediction, and the prediction result when the student model is not warmed up, i.e., K = 1, is the hard prediction. Calculate the total loss function of knowledge distillation, which is expressed as:

[0022] L = (1 - θ)L 1 + θL 2

[0023] wherein, L is the total loss function of knowledge distillation, L 1 is the distillation loss function, L 2 is the student loss function, θ is the weight coefficient; through repeated distillation training, optimize the structure and parameters of the student model, and finally obtain the optimal lightweight power vision large model.

[0024] Another object of the present invention is to provide a system for fine-tuning and lightweight design of a power vision large model, which can compress the fine-tuned power vision large model through a post-training optimization method based on quantization, optimize the parameters of the power vision large model through knowledge distillation, and output the optimal lightweight power vision large model, solving the problem of low quantization accuracy in the current model quantization technology.

[0025] As a preferred solution of the system for fine-tuning and lightweight design of the power vision large model described in the present invention, wherein: it includes a data processing module, a model fine-tuning module, and a model compression and optimization module; the data processing module is used to construct a power scene dataset based on the preprocessed power image data; the model fine-tuning module is used to fine-tune the power vision large model based on the power scene dataset and design a detection head; the model compression and optimization module is used to compress the fine-tuned power vision large model through a post-training optimization method based on quantization, optimize the parameters of the power vision large model through knowledge distillation, and output the optimal lightweight power vision large model.

[0026] A computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of a method for fine-tuning and lightweight design of a power vision large model.

[0027] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of a method for fine-tuning and lightweight design of a power vision large model.

[0028] Beneficial effects of the present invention: The method for fine-tuning and lightweight design of a large electric power vision model provided by the present invention is based on preprocessed electric power image data, and constructs and divides electric power scene data sets, which can effectively improve the generalization ability of the large electric power vision model in specific electric power scenes, and reduce the false alarm and missed alarm rates caused by single data distribution. By using visual sensors to collect electric power image data under various environmental conditions, and combining the preprocessing steps of cleaning, labeling and data enhancement, a high-quality electric power scene data set is successfully constructed and divided into a training set, a verification set and a test set. The data diversity is improved by multi-source data collection, the reliability and accuracy of the data are improved by image cleaning and labeling, and the coverage of data samples is expanded by data enhancement technology, and comprehensive modeling of electric power scenes is achieved, which can support complex deep learning tasks. The large electric power vision model is fine-tuned based on the electric power scene data set, and the specific adaptability of the large electric power vision model to the electric power scene is enhanced by a hierarchical fine-tuning strategy. The hierarchical fine-tuning strategy reduces development costs and improves efficiency, and realizes efficient fault identification and classification in complex electric power scenes. The hierarchical fine-tuning strategy avoids excessive adjustment of all parameters, which reduces the resources required for training pre-trained visual models. The invention can save time and effort, and improve the adaptability of the pre-trained visual model at a key level. The pre-trained visual model is initialized and the frozen layer is adjusted to adapt to the specific power scene characteristics. For data similar to the pre-trained model, some parameters are fine-tuned to reduce the waste of computing resources. For data with large differences, all parameters are fine-tuned, all network layers are unfrozen, and high-level classifiers and detectors are retrained. At the same time, a dedicated detection head is designed according to the target detection task, and the performance of the large power vision model is optimized by combining the classification loss function and the regression loss function. The network weights of the fine-tuned large power vision model are quantized by a quantitative post-training optimization method, and the floating-point weights are converted into a low-precision integer format, thereby realizing the compression of the large power vision model. The adaptive batch quantization of dynamic column combination solves the problem of uneven errors caused by the undifferentiation of column weight contributions in the traditional fixed order quantization method. Through contribution evaluation, dynamic column grouping and error feedback adjustment, the accumulation of quantization errors is reduced. Further combined with knowledge distillation, the teacher model and the student model are constructed, and the efficient migration of complex model knowledge is realized, which ensures the stability of the performance of the compressed large power vision model. The invention achieves better results in accuracy, adaptability and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0030] Figure 1 This is the overall flowchart of a method for fine-tuning and lightweight design of a power vision large model provided by the first embodiment of the present invention.

[0031] Figure 2 This is the schematic diagram of the optimization of the power vision large model based on knowledge distillation in a method for fine-tuning and lightweight design of a power vision large model provided by the first embodiment of the present invention.

[0032] Figure 3 This is the fine-tuning model diagram of the power professional model based on the vision large model in a method for fine-tuning and lightweight design of a power vision large model provided by the second embodiment of the present invention.

[0033] Figure 4 This is the schematic diagram of the modules of a system for fine-tuning and lightweight design of a power vision large model provided by the third embodiment of the present invention. Detailed implementation manners

[0034] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] Embodiment 1, referring to Figure 1 - Figure 2 , which is an embodiment of the present invention, provides a method for fine-tuning and lightweight design of a power vision large model, including:

[0036] S1: Based on the preprocessed power image data, construct a power scene dataset.

[0037] Furthermore, the preprocessed power image data includes collecting power image data through visual sensors at different times, locations, angles, lighting conditions, weather conditions, and seasons, including normal samples and abnormal samples of power equipment; the preprocessing includes cleaning the collected power image data, deleting blurred, duplicate, or irrelevant images, annotating the cleaned data, and further performing data augmentation processing on the annotated data.

[0038] It should be noted that constructing the power scene dataset includes dividing the power scene dataset into a training set, a validation set, and a test set. The training set is used to train the power vision large model, the validation set is used to determine the optimal power vision large model, and the performance of the optimal power vision large model is tested through the test set.

[0039] It should also be noted that power image data is collected through visual sensors, including but not limited to various visual sensors such as surveillance cameras, ball cameras, and drones. The samples are cleaned, specifically including methods such as image deduplication, abnormal image detection, and consistency detection. The cleaned data is labeled, including accurate labeling of the categories, positions, sizes, and other features of the objects in the images to ensure the accuracy and reliability of the data. After labeling, data cleaning is performed, including excluding blurred, duplicate, or irrelevant images to improve the quality and effectiveness of the dataset. The power scene dataset constructed after preprocessing will better support subsequent deep learning and application research.

[0040] S2: Fine-tune the power vision large model based on the power scene dataset.

[0041] Furthermore, fine-tuning the power vision large model includes fine-tuning the power vision large model based on the training set in the power scene dataset, adopting a hierarchical fine-tuning strategy, initializing the feature extractor of the pre-trained vision large model, loading the vision large model dedicated to the power industry, and fine-tuning the frozen layer selection parameters of the pre-trained vision large model feature extractor for fine-tuning or full-parameter fine-tuning. The feature extractor retains the low-level basic features of the pre-trained model and unfreezes the high-level parameters layer by layer, while retraining the high-level classifier and detector; when the data features and category distributions of the task scenario are similar to those of the pre-trained model, partial parameter fine-tuning is performed on the feature extractor, and only the high-level is unfrozen for fine-tuning; when the task scenario contains new device types, brand-new categories, or different fault features, which are quite different from the pre-trained model, full-parameter fine-tuning is performed to unfreeze all network layers.

[0042] It should be noted that fine-tuning the power vision large model also includes designing a detection head for the power scene. The detection head design selects a two-stage detector or a single-stage detector according to the object detection task, and constructs a total loss function through a classification loss function and a regression loss function; the cross-entropy loss is used to calculate the classification error and construct the classification loss function, which is expressed as:

[0043]

[0044] where L cls is the classification loss function, y i is the true label of category i, and p i is the predicted probability of category i by the power vision large model; the smooth L1 loss is used to calculate the bounding box regression error and construct the regression loss function, which is expressed as:

[0045]

[0046] where L regis the regression loss function, and x is the error between the true bounding box position t and the predicted position q; calculate the total loss function, expressed as:

[0047] L total = L cls + λL reg

[0048] where L total is the total loss function and λ is the balance factor.

[0049] It should also be noted that fine-tuning the large power vision model specifically includes initializing the feature extractor of the pre-trained large vision model, adjusting the frozen layer, and the design and integration of the detection head. The main purpose of initializing the feature extractor of the pre-trained large vision model is to complete the incremental pre-training in the power industry. This process requires loading a feature extractor of a large vision model specifically designed for the power industry, which can capture visual features in different scenarios. Adjusting the frozen layer is used to enhance the ability of the pre-trained large vision model to be applicable to different power scenarios. In different professional scenarios, according to the actual situation, all-parameter fine-tuning or partial-parameter fine-tuning of the feature extractor can be selected. Partial-parameter fine-tuning can maintain the powerful ability of the pre-trained model in feature extraction. By freezing some layers of the model and only unfreezing some high-level layers for fine-tuning, all-parameter fine-tuning is to unfreeze all layers of all networks. This method is more suitable for dealing with new tasks, especially when the dataset is large and has a large difference from the original pre-trained data. The design and integration of the detection head is to customize the detection head according to the specific object detection task. Common object detection frameworks include two-stage detectors (such as DETR) and single-stage detectors (such as EfficientDet). Each framework has its unique advantages and applicable scenarios. A reasonable detection head can improve the performance of the large power vision model.

[0050] S3: Compress the fine-tuned large power vision model based on the post-training optimization method of quantization, optimize the parameters of the large power vision model through knowledge distillation, and output the optimal lightweight large power vision model.

[0051] Furthermore, network weight quantization involves compressing the fine-tuned large-scale power vision model. Through the GPTQ quantization (post-training optimization of quantization) technique, the floating-point weights of the fine-tuned power vision model are converted into a low-precision integer format. The adoption of the GPTQ quantization technique includes adaptive batch quantization with dynamic column combination, and adaptive batch quantization with dynamic column combination includes contribution degree evaluation, dynamic column grouping, and error feedback adjustment. Contribution degree evaluation includes dividing each column vector of the weight matrix based on the LSH (Locality-Sensitive Hashing) technique and calculating the sensitivity index of each column weight in the model prediction accuracy. Dynamic column grouping includes grouping columns according to the sensitivity index and preferentially selecting column combinations with high contribution degrees and similar error distributions. Error feedback adjustment includes, after each batch of quantization is completed, re-evaluating the contribution degrees of the remaining columns based on the updated quantized column weights and dynamically adjusting the subsequent grouping strategy.

[0052] By adopting the GPTQ quantization technique including adaptive batch quantization with dynamic column combination, and adaptive batch quantization with dynamic column combination includes contribution degree evaluation, dynamic column grouping, and error feedback adjustment, network weight quantization is carried out. Different from traditional fixed-order batch quantization, adaptive batch quantization with dynamic column combination mainly solves the problem that the undifferentiated contribution degrees of column weights in the GPTQ method may lead to uneven errors, and optimizes the quantization effect by dynamically adjusting the column grouping order. The main advantages are as follows:

[0053] Improved quantization accuracy: By dynamically adjusting the column grouping order, the quantization error of high-contribution columns is minimized, while the error impact of low-contribution columns is balanced, reducing the possibility of overall quantization error accumulation.

[0054] Stable model performance: The dynamic column combination strategy effectively reduces the performance fluctuations introduced by weight quantization, especially showing high robustness in the batch quantization of large models.

[0055] High efficiency and adaptability: The adaptive batch quantization method with dynamic column combination does not rely on complex global optimization steps. Only the contribution degree evaluation and grouping adjustment modules are added in each batch, which neither changes the original batch quantization framework nor is applicable to neural networks of various scales.

[0056] Controllable computational overhead: Although the adaptive batch quantization method with dynamic column combination introduces dynamic adjustment steps, due to the efficient local sensitivity calculation of LSH and the hierarchical design of weight grouping, the overall computational complexity is only slightly higher than that of fixed-order batch quantization.

[0057] It should be noted that optimizing the parameters of the large - scale power vision model through knowledge distillation and outputting the optimal lightweight large - scale power vision model includes constructing a teacher model and a student model. The fine - tuned and compressed large - scale power vision model is used as the teacher model, and the soft labels generated by the teacher model are passed to the student model. By introducing a distillation temperature to smooth the prediction results, it is expressed as:

[0058]

[0059] Among them, y i is the prediction result of class i, z i is the original score of class i, K is the distillation temperature, zj is the original score of all classes j; the prediction result after the teacher model is warmed up is the soft label, and the prediction result after the student model is warmed up is the soft prediction. When the student model is not warmed up, that is, K = 1, the prediction result is the hard prediction. Calculate the total loss function of knowledge distillation, which is expressed as:

[0060] L=(1 - α)L 1 +αL 2

[0061] Among them, L is the total loss function of knowledge distillation, L 1 is the distillation loss function, L 2 is the student loss function, and α is the weight coefficient; through repeated distillation training, optimize the structure and parameters of the student model, and finally obtain the optimal lightweight large - scale power vision model.

[0062] It should also be noted that network weight quantization of the fine - tuned large - scale power vision model based on the GPTQ method is to perform weight quantization processing on the already trained large - scale power vision model. By analyzing the distribution of the model weights, determine the appropriate quantization bits, and use the quantization algorithm to convert the continuous weight values into discrete quantization values, so as to reduce the storage space and computational burden of the model, and then improve the operation efficiency of the model. The GPTQ method can be used to quantize the weights of the fine - tuned large - scale power vision model by using the AutoGPTQ tool. The specific steps are as follows:

[0063] First, it is necessary to prepare the relevant development environment, install the necessary libraries through pip, such as PyTorch and Transformers, etc. Next, ensure that the used large-scale power vision model is well-trained and has good performance to lay the foundation for subsequent quantization. After loading the model, the relevant parameters for quantization need to be configured next. These parameters include the target bit width, whether to use per-channel quantization, etc., to improve the flexibility and effect of quantization. By adjusting these settings according to specific requirements, the advantages of quantization can be maximized to ensure that the model can still maintain good inference performance after quantization. After the configuration is completed, use the interface of AutoGPTQ to apply the GPTQ quantization method. This process will process the weights of the model and convert the floating-point weights into fixed-point format, thus effectively reducing the volume and memory occupancy of the model. After quantization is completed, save the quantized model locally for quick loading and inference. Finally, by loading the saved quantized model, the inference speed can be improved while reducing resource consumption.

[0064] It should also be noted that optimizing the parameters of the large-scale power vision model through knowledge distillation uses the large-scale power vision model after constructing a power scene dataset, fine-tuning and compressing the large-scale power vision model as the teacher model, constructing a more concise student model, and realizing the knowledge transfer process from the complex teacher model to the simple student model. As Figure 2 shown, in the knowledge distillation technology, in order to make the probability distribution smoother and reduce the overfitting risk of the model, on the basis of the original softmax formula, the output of the softmax layer is improved by introducing a new hyperparameter distillation temperature K, so as to obtain the final prediction result of each category. The loss function in the training process consists of two parts. One is the distillation loss, that is, the difference between the soft prediction of the student model and the soft label of the teacher model at the distillation temperature K. The other is the student loss, that is, the difference between the hard prediction of the student model and the real result without using the distillation temperature. Through the knowledge distillation technology, the complexity of the original teacher model can be effectively reduced, and at the same time, it is ensured that the new student model can maintain a high prediction accuracy. In this way, not only can the operation efficiency of the large-scale power vision model be improved, but also its performance in practical applications can be ensured to meet the expected standards, so as to better serve the safety monitoring and management of the power system.

[0065] Example 2, referring to Figure 3 , which is an embodiment of the present invention, provides a method for fine-tuning and lightweight design of a large-scale power vision model. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0066] To verify the optimization effect of the present invention on the power vision large model, specifically, there are currently 30,000 image samples of air switch on, air switch off, construction workers not wearing safety helmets, construction workers not wearing work uniforms, meter blurring, and damaged marker shells collected by devices such as cloth control cameras, drones, and surveillance cameras. Among them, there are 5,000 selected samples for each category, 4,000 are used for model training, and 1,000 are used for model inference. In total, there are 24,000 training sets and 6,000 test sets. Now, the power vision large model (L1), the fine-tuning model (L2) of the power professional model based on the vision large model as Figure 3 shown, the compression model (L3) that quantizes the network weights based on the L2 model, and the knowledge distillation model (L4) based on L3 are respectively tested.

[0067] Referring to Table 1 and Table 2, the experimental data is analyzed and compared.

[0068] Table 1 Test results at each stage

[0069]

[0070]

[0071] Table 2 Model inference time consumption and video memory occupancy at each stage

[0072] Model L1 L2 L3 L4 Inference Time per Image 400ms 400ms 150ms 50ms Video Memory Occupancy 8GB 8GB 6GB 2.7GB

[0073] The test results are shown in Table 1 and Table 2. According to the analysis of the experimental results, after a series of model optimizations such as model fine-tuning, network weight quantization compression, and knowledge distillation for 6 scenarios, the overall accuracy of the L4 model is improved by about 10%, the recall rate is improved by about 8%, and the AP is improved by about 7%. AP is an index that combines precision and recall rate; while the accuracy of the L4 model is improved, the inference time consumption is only 12.5% of the original model, and the video memory occupancy is only 33.7% of the original model, greatly reducing the model deployment cost. Therefore, the present invention has creativity.

[0074] Example 3, referring to Figure 4 , is an embodiment of the present invention, which provides a system for fine-tuning and lightweight design of a power vision large model, including a data processing module, a model fine-tuning module, and a model compression optimization module.

[0075] Among them, the data processing module is used to construct a power scenario dataset based on the preprocessed power image data; the model fine-tuning module is used to fine-tune the power vision large model based on the power scenario dataset and design a detection head; the model compression optimization module is used to perform network weight quantization on the fine-tuned power vision large model based on the quantization post-training optimization method, optimize the parameters of the power vision large model through knowledge distillation, and output the optimal lightweight power vision large model.

[0076] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., all kinds of media that can store program codes.

[0077] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0078] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROMs). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.

[0079] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

[0080] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for fine-tuning and lightweight design of a large electric visual model, characterized in that: include: Based on the preprocessed power image data, a power scene dataset is constructed; Fine-tune the large power vision model based on the power scene dataset; The fine-tuned power vision big model is compressed based on the quantization-based post-training optimization method, the parameters of the power vision big model are optimized through knowledge distillation, and the optimal lightweight power vision big model is output.

2. The method for fine-tuning and lightweight design of a large electric visual model as claimed in claim 1, characterized in that: The pre-processed power image data includes collecting power image data, including normal samples and abnormal samples of power equipment, at different times, locations, angles, lighting conditions, weather conditions and seasons through visual sensors; Preprocessing includes cleaning the collected power image data, deleting fuzzy, repeated or irrelevant images, labeling the cleaned data, and further using data enhancement processing on the labeled data.

3. The method for fine-tuning and lightweight design of a large electric visual model as claimed in claim 2, characterized in that: The construction of the power scene data set includes dividing the power scene data set into a training set, a validation set and a test set, the training set is used to train the power vision big model, the validation set is used to determine the optimal power vision big model, and the test set is used to test the performance of the optimal power vision big model.

4. The method for fine-tuning and lightweight design of a large electric visual model as claimed in claim 3, characterized in that: The fine-tuning of the electric power vision big model includes fine-tuning the electric power vision big model based on the training set in the electric power scene data set, adopting a layered fine-tuning strategy, initializing the pre-trained visual big model feature extractor, loading the electric power industry-specific visual big model, and adjusting the pre-trained visual big model feature extractor to freeze the layer to select parameter fine-tuning or full parameter fine-tuning. The feature extractor retains the low-level basic features of the pre-trained model, and unfreezes the high-level parameters layer by layer, and retrains the high-level classifier and detector at the same time; When the data features and category distribution of the task scenario are similar to those of the pre-trained model, fine-tune some parameters of the feature extractor and only unfreeze the high-level layers for fine-tuning; When the task scenario contains new equipment types, completely new categories, or different fault characteristics that are significantly different from the pre-trained model, all parameters are fine-tuned and all network layers are unfrozen.

5. The method for fine-tuning and lightweight design of a large electric visual model as claimed in claim 4, characterized in that: Also includes: Designing a detection head, wherein the designing of the detection head includes designing a detection head for a power scenario, wherein the detection head design selects a two-stage detector or a single-stage detector according to a target detection task, and constructs a total loss function through a classification loss function and a regression loss function; The cross entropy loss is used to calculate the classification error and construct the classification loss function, which is expressed as: Among them, L cls is the classification loss function, y i is the true label of category i, p i is the predicted probability of category i by the power vision model; The smooth L1 loss is used to calculate the bounding box regression error and construct the regression loss function, which is expressed as: Among them, L reg is the regression loss function, x is the error between the true bounding box position t and the predicted position q; Calculate the total loss function, expressed as: THE total =L cls +λL reg Among them, L total is the total loss function, and λ is the balancing factor.

6. The method for fine-tuning and lightweight design of a large electric visual model as claimed in claim 5, characterized in that: The compression includes quantizing the network weights of the fine-tuned power vision model, and converting the floating-point weights of the fine-tuned power vision model into a low-precision integer format through a quantized post-training optimization technique; Post-training optimization techniques using quantization include adaptive batch quantization of dynamic column combinations, which includes contribution evaluation, dynamic column grouping, and error feedback adjustment; The contribution evaluation includes partitioning each column vector of the weight matrix based on the local sensitive hashing technique and calculating the sensitivity index of each column weight in the model prediction accuracy; Dynamic column grouping includes grouping columns according to sensitivity indicators, giving priority to column combinations with high contribution and similar error distribution; Error feedback adjustment involves re-evaluating the contribution of the remaining columns based on the updated quantization column weights after the quantization of each batch is completed, and dynamically adjusting the subsequent grouping strategy.

7. The method for fine-tuning and lightweight design of a large electric visual model as claimed in claim 6, characterized in that: The method of optimizing the parameters of the electric power vision big model through knowledge distillation and outputting the optimal lightweight electric power vision big model includes constructing a teacher model and a student model, using the fine-tuned and compressed electric power vision big model as the teacher model, and passing the soft label generated by the teacher model to the student model, and introducing the distillation temperature to smooth the prediction result, which is expressed as: Among them, y i is the prediction result of category i, z i is the original score of category i, K is the distillation temperature, z j is the raw score of all categories j; The prediction result after the teacher model is warmed up is a soft label, the prediction result after the student model is warmed up is a soft prediction, and the prediction result when the student model is not warmed up, that is, K = 1, is a hard prediction. The total loss function of knowledge distillation is calculated and expressed as: L=(1-α)L1+αL2 Among them, L is the total loss function of knowledge distillation, L1 is the distillation loss function, L2 is the student loss function, and α is the weight coefficient; Through repeated distillation training, the structure and parameters of the student model are optimized, and finally the optimal lightweight power vision large model is obtained.

8. A system using the method for fine-tuning and lightweight design of a large electric visual model as described in any one of claims 1 to 7, characterized in that: Including data processing module, model fine-tuning module, model compression and optimization module; The data processing module is used to construct a power scene data set based on the preprocessed power image data; The model fine-tuning module is used to fine-tune the large power vision model based on the power scene data set and design a detection head; The model compression optimization module is used to compress the fine-tuned power vision large model based on a quantized post-training optimization method, optimize the power vision large model parameters through knowledge distillation, and output the optimal lightweight power vision large model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for fine-tuning and lightweight design of a large electric visual model according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for fine-tuning and lightweight design of a large electric visual model according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Large-model cross-end transfer compression method and device applied to intelligent electric energy meter

    CN120321130A

  • Coal mine drilling center positioning method based on image processing and deep learning

    CN120707627A

  • Small sample learning optimization method and system for image classification, medium and equipment

    CN120726373A