Lightweight deployment model edge configuration method based on damper multi-defect detection
By performing scene similarity processing on the anti-vibration hammer multi-defect data set and edge device attribute data set, sparse training and combining knowledge distillation strategies, a lightweight model with strong adaptability is built, which solves the problems of large amount of calculation and poor adaptability of deep learning models when deploying on edge devices, and achieves efficient anti-vibration hammer defect detection.
Patent Information
- Application Number
- CN202510176810.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-11
AI Technical Summary
The existing deep learning models are large in computing and poor in adaptability when deployed on edge devices, making it difficult to adapt to different defect scenarios and edge device resource characteristics, resulting in low detection efficiency and waste of resources.
By processing scene similarity of the anti-vibration hammer multi-defect data set and edge device attribute data set, building a target scene sequence, sparsely training the initial teacher model, and adjusting the initial student model through knowledge distillation strategies to adapt to different defect scenarios and edge device resource characteristics, achieving lightweight deployment.
The accuracy and efficiency of anti-vibration hammer defect detection are improved, the problem of poor adaptability of the model on edge devices is solved, and efficient deployment and optimized resource utilization are achieved.
Smart Images

Figure CN120298850A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent power grid operation and maintenance. Specifically, it relates to a lightweight deployment model edge configuration method based on multi-defect detection of vibration dampers. Background Art
[0002] In the stable operation of transmission lines, vibration dampers play a crucial role. Once a vibration damper has defects, such as component loosening, cracks, etc., it will greatly affect the safety of transmission lines and even cause serious accidents. The traditional manual inspection method is inefficient, costly, and difficult to achieve real-time monitoring. With the development of artificial intelligence technology, using deep learning models for multi-defect detection of vibration dampers has become a research hotspot. However, deep learning models usually have complex structures and large computational amounts, and directly deploying them on edge devices faces many challenges. For example, the computing power, storage capacity, and network bandwidth of edge devices are limited, making it difficult to support the efficient operation of complex models; existing model deployment methods lack effective consideration of the actual on-site scenarios and do not fully combine the vibration damper defect scenarios and the resource characteristics of edge devices for targeted configuration, resulting in poor adaptability of the model to the actual scenarios and being unable to fully exert the performance of the model; at the same time, when dealing with multi-defect detection, the adaptability to different defect types and scenarios is poor.
[0003] Chinese Patent, Publication No.: CN 119228767 A, Publication Date: December 31, 2024, discloses a lightweight transmission line insulator defect detection method based on knowledge distillation. The detection method is as follows: constructing an insulator defect detection model based on knowledge distillation, and inputting the real-time collected insulator image data into the insulator defect detection model based on knowledge distillation to obtain the insulator defect detection result; the invention trains a smaller student network by transferring knowledge from a larger teacher model through a knowledge distillation module, which can improve the accuracy of insulator defect recognition and enable it to be deployed on devices with lower computing performance, having strong practicability. Since the solution does not consider the complex power grid topology in the power operation site and the actual situation that the performance of distributed edge computing devices varies, it is difficult to obtain a student model that adapts to the performance requirements of different edge devices, thus greatly hindering the efficiency of scenario fault detection and causing a great waste of edge computing resources.
[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present application. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the present invention proposes a lightweight deployment model edge configuration method based on vibration damper multi-defect detection. By performing scene similarity processing on the vibration damper multi-defect data set and the edge device attribute data set, a target scene sequence is obtained. On this basis, the initial teacher model is sparsely trained, and the parameters of the initial student model are adjusted through a knowledge distillation strategy, so that the obtained target student model can better adapt to different defect scenarios and edge device resource characteristics, thereby improving the accuracy and efficiency of vibration damper defect detection, and at the same time realizing the efficient deployment of the model on edge devices.
[0006] A technical solution provided in an embodiment of the present invention is a lightweight deployment model edge configuration method based on vibration damper multi-defect detection, including the following steps: S1. Perform scene similarity processing on the vibration damper multi-defect data set and the edge device attribute data set to obtain a target scene sequence; S2. Sparsely train the initial teacher model in turn through the target scene sequence to obtain a target teacher model and a corresponding initial weight sequence; construct an initial student model based on the initial weight sequence; S3. Adjust the parameters of the initial student model through a knowledge distillation strategy in combination with the target teacher model to obtain a target student model; S4. The target student model adapts to the target scene sequence to obtain an edge configuration path, and implant the target student model into the corresponding edge device according to the edge configuration path.
[0007] Preferably, the performing scene similarity processing on the vibration damper multi-defect data set and the edge device attribute data set to obtain a target scene sequence includes the following steps: S11. Obtain vibration damper defect images in the power transmission management domain to construct a vibration damper multi-defect data set and all edge device information to construct an edge device attribute data set; S12. Cluster the elements in the vibration damper multi-defect data set according to defect similarity to obtain a defect sequence; Synchronously, cluster the elements in the edge device attribute data set according to resource similarity to obtain a resource sequence; S13. Extract the edge device information corresponding to the defect scenario and integrate the defect sequence and the resource sequence to obtain a target scene sequence.
[0008] Preferably, the S12 includes the following steps: Extract the color features, texture features, and shape features of the vibration damper defect images, calculate the cosine similarity between all images in the vibration damper multi-defect data set through the cosine similarity algorithm, and divide the images with a cosine similarity higher than the set threshold into the same class, thereby obtaining a defect sequence.
[0009] Preferably, the S12 further includes the following steps: Extract the computing power characteristics, storage capacity characteristics, and network bandwidth characteristics of edge devices. Calculate the Euclidean distance between all pairs of edge devices in the edge device attribute dataset through the Euclidean distance algorithm. Divide the edge devices with Euclidean distance less than the set threshold into the same class, thereby obtaining a resource sequence.
[0010] Preferably, the method for integrating the defect sequence and the resource sequence by using the edge device information corresponding to the defect scenario to obtain a target scenario sequence includes the following steps: S131. Search for all edge devices within the area centered on the defect scenario with a search radius of r. S132. Calculate the scenario adaptation degree according to the connection degree, communication distance, and resource matching degree between the edge device and the defect scenario; S133. Sort the edge devices according to the scenario adaptation degree to determine the target scenario sequence corresponding to each defect scenario.
[0011] Preferably, sparsely train the initial teacher model in sequence through the target scenario sequence to obtain a target teacher model and a corresponding initial weight sequence; construct an initial student model based on the initial weight sequence; the method includes the following steps: S21. Train the initial teacher model through the target scenario sequence to obtain a batch dataset, normalize the batch dataset, and correct the scaling factor through the preset scaling factor and bias parameter of the channel to obtain a target scaling factor; S22. Adopt the L1 regularization method to sparsely express the set Γ of target scaling factors corresponding to each channel in the convolutional layer; connect the sparsely trained target scaling factors with the corresponding channels in the convolutional layer feature map, and sort them by taking the absolute value to obtain a joint sequence; S23. Determine the pruning threshold according to the required computing power value of the basic model and the actual maximum computing power value of the edge device; if the target scaling factor corresponding to the feature channel is less than the pruning threshold, determine it as a secondary channel and apply pruning to the joint sequence to obtain a target teacher model; S24. Record the weight change after each pruning to obtain an initial weight sequence; use the initial weight sequence as the training weight of the student model to construct an initial student model.
[0012] The method for adjusting the parameters of the initial student model by using the knowledge distillation strategy in combination with the target teacher model to obtain a target student model includes the following steps: S31. Initialize the target teacher model and the initial student model so that their channel numbers and feature map sizes are adapted; use the KL divergence to measure the feature difference between the target teacher model and the initial student model on each feature map to obtain an internal hierarchical feature distillation loss value; S32, calculating the difference between the bounding box predicted by the target teacher model and the bounding box predicted by the initial teacher model based on the shape intersection-over-union metric to obtain the bounding box distillation loss; Simultaneously, the classification distillation loss is determined by measuring the difference between the classification predictions of the initial student model and the classification predictions of the target teacher model; S33. Calculate the external level feature distillation loss value through the bounding box distillation loss and the classification distillation loss; obtain the fused feature distillation loss by weighted summing the internal level feature distillation loss value, the external level feature distillation loss value and their corresponding weight values; solve the objective function with the minimum fused feature distillation loss to obtain the parameter adjustment value, and obtain the target student model by correcting the parameter adjustment value in the initial student model.
[0013] Preferably, the method of using KL divergence to measure the feature difference between the target teacher model and the initial student model on each feature map to obtain the internal level feature distillation loss value includes the following steps: Calculate the Euclidean distance between the feature vector and all other feature vectors in the batch to obtain a distance matrix. Create a mask matrix with the same dimension as the distance matrix based on the category label to identify samples with the same label. Find the maximum distance among samples with the same label as the hardest positive sample distance, and the minimum distance among samples with different labels as the hardest negative sample distance; The distance loss value between the hardest positive sample and the hardest negative sample is calculated through the triplet loss; the feature difference is measured by the distance loss to obtain the bounding box distillation loss.
[0014] Preferably, the step of measuring the difference between the classification prediction of the initial student model and the classification prediction of the target teacher model to determine the classification distillation loss comprises the following steps: The classification logits of the target teacher model and the initial student model are normalized by sigmoid, and the binary cross entropy loss between the classification prediction of the initial student model and the classification prediction of the teacher model is calculated to obtain the classification prediction difference to determine the classification distillation loss.
[0015] Preferably, the target student model is adapted to the target scene sequence to obtain an edge configuration path, and the target student model is implanted into a corresponding edge device according to the edge configuration path, including the following steps: S41, setting a scene adaptability threshold, screening out target scene sequences greater than the scene adaptability threshold, and calculating the rated working computing power required for each fault scene; determining edge devices with values greater than the rated working computing power as pre-selected edge devices; S42. Determine the model computing power value of the target student model according to the distillation loss. Use the difference between the model computing power value and the rated working computing power value as the numerator, and use the difference between the actual maximum computing power value of the rated working computing power value and the rated working computing power value as the denominator; obtain the computing power resource utilization rate, and use the preselected edge device with the maximum computing power resource utilization rate as the target edge device; S43. Construct a model receiving channel for the target edge device through an asymmetric encryption algorithm to obtain an edge configuration path, and the target student model is sent to the corresponding target edge device through the edge configuration path.
[0016] Advantages of the present invention: (1) By performing scene similarity processing on the anti-vibration hammer multi-defect dataset and the edge device attribute dataset, the present invention obtains a target scene sequence. On this basis, the initial teacher model is sparsely trained, and the parameters of the initial student model are adjusted through a knowledge distillation strategy, so that the model can better adapt to different defect scenarios and edge device resource characteristics, thereby improving the accuracy and efficiency of anti-vibration hammer defect detection, and avoiding the problems of low efficiency, high cost and difficult real-time monitoring of traditional manual inspections; (2) Aiming at the problems of limited computing power, storage capacity and network bandwidth of edge devices, the present invention makes targeted configurations in combination with the anti-vibration hammer defect scenario and the resource characteristics of edge devices. By sparsely training the initial teacher model, the pruning threshold is determined according to the computing power value required by the basic model and the actual maximum computing power value of the edge device to obtain the target teacher model, and the initial student model is constructed. Finally, the target student model is adapted to the target scene sequence to obtain the edge configuration path, realizing the efficient deployment of the model on the edge device and solving the problem of poor adaptability caused by the lack of effective consideration of the actual scene on site in the existing model deployment method; (3) When dealing with multi-defect detection, the present invention clusters the anti-vibration hammer multi-defect dataset to obtain a defect sequence, comprehensively considers factors such as the connection degree, communication distance and resource matching degree between the edge device and the defect scenario to determine the scene adaptability, and then ranks the edge devices by priority to determine the target scene sequence. At the same time, in the process of model training, through a variety of loss calculation methods, such as internal hierarchical feature distillation loss value, bounding box distillation loss, classification distillation loss, etc., and fusing these loss values to adjust the model parameters, so that the model has better adaptability to different defect types and scenarios, overcoming the problem of poor adaptability in multi-defect detection in the prior art.
[0017] The above description of the invention content is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically described below. Description of the Drawings
[0018] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments read in conjunction with the accompanying drawings. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0019] Figure 1 It is a flowchart of the edge configuration method of the lightweight deployment model based on vibration damper multi-defect detection of the present invention.
[0020] Figure 2 It is the overall training network framework diagram of this embodiment.
[0021] Figure 3 It is the structural framework diagram of the ILFD module of this embodiment.
[0022] Figure 4 It is the structural framework diagram of the OLKD module of this embodiment.
[0023] Figure 5 It is the flowchart of the target scene sequence generation of the embodiment of the present invention.
[0024] Figure 6 It is the flowchart of the initial student model construction of the embodiment of the present invention. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific implementation manners described herein are only the best embodiments of the present invention, which are only used to explain the present invention and do not limit the protection scope of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0026] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the operations (or steps) as sequential processes, many of the operations (or steps) can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings; the process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0027] Embodiment: As Figure 1 shown, the edge configuration method of the lightweight deployment model based on vibration damper multi-defect detection includes the following steps: S1. Perform scene similarity processing on the anti-vibration hammer multi-defect dataset and the edge device attribute dataset to obtain the target scene sequence.
[0028] Specifically, as Figure 5 shown, S1 includes the following steps: S11. Obtain the anti-vibration hammer defect images within the power transmission management domain to construct the anti-vibration hammer multi-defect dataset, and obtain all edge device information to construct the edge device attribute dataset.
[0029] It can be understood that, assuming in a specific power transmission management domain, this area contains multiple power transmission line segments, and different models of anti-vibration hammers are installed on each segment. Use a drone equipped with a high-definition camera to regularly take pictures of the anti-vibration hammers on the power transmission line to obtain a large number of anti-vibration hammer images; in order to obtain rich anti-vibration hammer images, a dataset of anti-vibration hammer defect images downloaded from the public network can be obtained through web crawler technology. Screen the obtained images, remove invalid images such as blurry ones and those with poor shooting angles, and label the clear images containing defect information. The labeling content includes defect types (such as component looseness, cracks, deformation, etc.), defect locations, etc., so as to construct the anti-vibration hammer multi-defect dataset. At the same time, collect the information of all edge devices within this power transmission management domain, including device models, computing capabilities (such as the number of CPU cores, main frequency, etc.), storage capacities (memory size, storage hard disk capacity), network bandwidths (upstream and downstream bandwidths), device locations (latitude and longitude information), etc., and thus construct the edge device attribute dataset.
[0030] S12. Cluster the elements in the anti-vibration hammer multi-defect dataset according to defect similarity to obtain the defect sequence; Synchronously, cluster the elements in the edge device attribute dataset according to resource similarity to obtain the resource sequence.
[0031] As an optional embodiment, S12 further includes the following steps: Extract the color features, texture features, and shape features of the anti-vibration hammer defect images, calculate the cosine similarity between all images in the anti-vibration hammer multi-defect dataset through the cosine similarity algorithm, and divide the images with a cosine similarity higher than the set threshold into the same class, thereby obtaining the defect sequence.
[0032] It is understandable that in the multi-defect dataset of vibration damping hammers, an image feature extraction algorithm, such as a feature extractor based on a convolutional neural network, is used to extract the color features (such as color differences in different defect parts), texture features (such as texture details of cracks), and shape features (shape changes after component deformation) of the vibration damping hammer defect images. Then, the cosine similarity algorithm is used to calculate the cosine similarity between all pairs of images in the dataset. A suitable threshold is set, for example, 0.8, and images with a cosine similarity higher than this threshold are grouped into the same class to obtain a defect sequence. For example, all images of component loosening are grouped into one class, and crack images are grouped into another class, etc. For example, assume that the feature vectors of two vibration damping hammer defect images are A = [a1, a2, ···, a n and B = [b1, b2, ···, b n , where c j and d j respectively represent different defect feature images of the device, and these feature vectors can be obtained by extracting features such as color, shape, and texture of the image through algorithms such as Scale-Invariant Feature Transform (SIFT); exemplarily, the cosine similarity calculation formula is:
[0033] As an alternative embodiment, S12 further includes the following steps: Extract the computing power characteristics, storage capacity characteristics, and network bandwidth characteristics of the edge device, calculate the Euclidean distance between all pairs of edge devices in the edge device attribute dataset through the Euclidean distance algorithm, and group edge devices with a Euclidean distance less than the set threshold into the same class to obtain a resource sequence.
[0034] It is understandable that for the edge device attribute dataset, the computing power characteristics, storage capacity characteristics, and network bandwidth characteristics of the edge device are extracted. Using the Euclidean distance algorithm, the Euclidean distance between all pairs of edge devices in the dataset is calculated. A distance threshold is set, such as edge devices with a distance less than a certain value (determined according to the actual situation) are grouped into the same class to obtain a resource sequence. For example, edge devices with similar computing power, similar storage capacity, and the same network bandwidth level are grouped into one class. In terms of resource similarity measurement, the Euclidean distance algorithm is used. Let the resource attribute vectors of two edge devices be C = [c1, c2, ···, c m and D = [d1, d2, ···, d m , where c j and d j respectively represent different resource attributes of the device, such as attribute values of computing power, storage capacity, network bandwidth, etc., and the Euclidean distance calculation formula is:
[0035] S13. Extract the edge device information corresponding to the defect scenario, and integrate the defect sequence and the resource sequence to obtain the target scenario sequence.
[0036] As an alternative embodiment, S13 includes the following steps: S131. Search for all edge devices within a region centered on the defect scenario with a search radius of r. S132. Calculate the scenario fitness based on the connection degree, communication distance, and resource matching degree between the edge device and the defect scenario; S133. Sort the edge devices according to the scenario fitness to determine the target scenario sequence corresponding to each defect scenario.
[0037] It can be understood that with the help of a high-precision geolocation system and combined with the geographical information data of the transmission line, the geographical position coordinates of the defect scenario are accurately determined. The value of the radius r is comprehensively determined based on the actual distribution density of the transmission line, the coverage range of the edge device, and the data transmission delay requirements. For example, in a densely populated area of the transmission line, r can be set to 2 - 3 kilometers to ensure that a sufficient number of edge devices can be covered; while in a relatively sparse area, r can be appropriately expanded to 5 - 8 kilometers. Using the spatial query function of the Geographic Information System (GIS), the search area is circled on the map according to the set radius r, and information such as the unique identifier, geographical position, and device status of all edge devices within this area is retrieved to establish a preliminary candidate set of edge devices. Perform basic condition screening on the preliminarily retrieved edge devices. According to the pre-set basic resource requirements, such as the computing power not less than [Z] and the storage capacity not less than [W], check the hardware parameters of the edge devices. By querying the edge device attribute dataset, obtain the computing power (such as the number of CPU cores, main frequency, etc.) and storage capacity (memory, hard disk capacity, etc.) information of each device, exclude the devices that do not meet the requirements, narrow the device search range, and improve the subsequent processing efficiency.
[0038] Furthermore, the connection degree is used to measure the tightness of the edge device and the defect scenario in terms of network topology or physical connection. If the edge device and the location of the defect scenario are in the same subnet or directly connected network nodes, the connection degree is 1; if it passes through 1 - 2 network node transfers in the middle, the connection degree is 0.8; if it passes through 3 - 4 network node transfers, the connection degree is 0.6; and so on. After passing through g network node transfers, the connection degree H is 1 - 0.2(g - 1) (g is an integer greater than 1). Through the network topology analysis tool, obtain the network connection path between the edge device and the defect scenario, determine the number of transfer network nodes, and thus calculate the connection degree.
[0039] Furthermore, if the longitude and latitude information of the device is known, the great circle distance formula can be used where r is the radius of the earth, and Let $\varphi$ be the latitude of two positions and $\Delta\lambda$ be the longitude difference. Calculate the geographical distance between the edge device and the defect scenario, and then convert the geographical distance into a communication distance according to the relationship between the actual network transmission loss and the distance. If only the network topology relationship is known, the communication distance can be estimated by the network hop count, with each hop set as a unit distance. The calculated communication distance needs to be normalized so that its value is between 0 and 1 for comprehensive calculation of the scenario fitness with other factors.
[0040] Furthermore, establish a resource matching model, comprehensively considering factors such as computing power, storage capacity, and network bandwidth. Assume that the weight of computing power is $w_1$, the weight of storage capacity is $w_2$, and the weight of network bandwidth is $w_3$, where $w_1 + w_2 + w_3 = 1$. For computing power, compare the actual computing power of the edge device with the computing power threshold required to process the defect scenario to obtain the computing power matching score $S_1$; similarly, obtain the storage capacity matching score $S_2$ and the network bandwidth matching score $S_3$. The resource matching degree $R = w_1S_1 + w_2S_2 + w_3S_3$. Therefore, the scenario fitness $F = H\cdot(1 - D)\cdot R$, where $H$ is the connectivity, $D$ is the normalized communication distance, and $R$ is the resource matching degree. The formula comprehensively considers the connectivity, communication distance, and resource matching degree. The higher the connectivity, the closer the communication distance, and the better the resource matching degree, the higher the scenario fitness.
[0041] It can be understood that according to the calculated scenario fitness, the edge devices are sorted in descending order. The edge device with the highest scenario fitness is given the highest priority, and each defect scenario is associated with the edge devices with decreasing priority in turn to determine the initial target scenario sequence. During the association process, record the defect scenarios assigned to each edge device and the corresponding scenario fitness values. When a high-priority device fails, is occupied by other defect scenarios, or compatibility issues between certain devices and the model are found during actual deployment testing, automatically remove the device from the association list of the current defect scenario, recalculate the priorities of the remaining devices, and select the second-priority device for association. At the same time, re-verify the adjusted target scenario sequence to ensure that each defect scenario can find a suitable edge device and the overall scenario fitness is optimal. With the operation and maintenance of the transmission line, new defect scenarios may occur for the vibration dampers, and the states and resources of the edge devices may also change. Recalculate the scenario fitness between each edge device and the new defect scenario or the changed scenario, re-perform the priority sorting and association to ensure the accuracy and effectiveness of the target scenario sequence to adapt to the continuously changing actual situation.
[0042] S2. Sparsely train the initial teacher model through the target scenario sequence in turn to obtain the target teacher model and the corresponding initial weight sequence; construct the initial student model based on the initial weight sequence.
[0043] Specifically, asFigure 6 As shown in Figure 6 , S2 includes the following steps: S21. Train the initial teacher model through the target scenario sequence to obtain a batch dataset, normalize the batch dataset, and correct the scaling factor through the preset scaling factor and bias parameter of the channel to obtain the target scaling factor; S22. Adopt the L1 regularization method to sparsely represent the set Γ of the target scaling factors corresponding to each channel in the convolutional layer; connect the sparsely trained target scaling factors with the corresponding channels in the convolutional layer feature map, and sort them by taking the absolute value to obtain the joint sequence; S23. Determine the pruning threshold according to the required computing power value of the basic model and the actual maximum computing power value of the edge device; if the target scaling factor corresponding to the feature channel is less than the pruning threshold, determine it as a secondary channel and apply pruning to the joint sequence to obtain the target teacher model; S24. Record the weight change after each pruning to obtain the initial weight sequence; use the initial weight sequence as the training weight of the student model to construct the initial student model.
[0044] In this embodiment, due to the limited computing power of the memory of the embedded edge device, it is challenging in the multi-defect detection task of vibration dampers. Therefore, this application adopts the pruning scheme shown in the lower half to prune the trained basic model without changing the existing CNN architecture. First, assign a scaling factor to each channel of the trained basic model to represent the importance of the channel. Then, perform channel sparse training to distinguish important channels from unimportant channels. Finally, set a global threshold to determine the channels to be pruned to obtain the final pruned student model. Figure 2 Figure 2
[0045] Furthermore, the following example will illustrate the process of scaling factor selection and sparse training: The purpose of sparse training is to discriminate the importance of each channel in the convolutional layer feature map and provide a reference index for channel pruning. To avoid introducing additional computational overhead, the scaling factor γ in the BN (normalization layer) is sparsely processed as the basis for judging channel importance. Let the output of the previous layer be x1, x2,..., x m , m is the number of training sample batches, μ B and σ B are the mean and variance of each batch, ε is a regularization parameter used to prevent the denominator from being zero during calculation, and the normalized output is as shown in Equation (1). To improve the non-linear feature extraction ability and overall expression ability of the model, the scaling factor is reconstructed through the learnable scaling factor γ and bias parameter β, as shown in Equation (2), where y i is the reconstruction result; Adopt the L1 regularization method to sparsely represent the set Γ of γ corresponding to each channel in the convolutional layer. After sparsification, the feature channels where γ tends to 0 are the secondary channels. Introduce the regularization formula of Γ into the network loss function Loss: where l(f(x, W), y) is the loss function of the original convolutional neural network, (x, y) are the training inputs and targets, W is the trainable weight, λ is the balance factor used to control different degrees of channel sparse regularization, and g(γ) is the sparse penalty term for the scaling factor, that is, g(γ) = ||γ||1.
[0046] Based on the sparsely trained scaling factor γ, connect it to the corresponding channels in the convolutional layer feature map, and sort by taking the absolute value to obtain a joint sequence. As shown in Equation (4), determine the pruning threshold T by setting the pruning ratio Pratio, GFLOPsbase is the required computing power value of the base model, and GFLOPspruned is the actual maximum computing power value of the edge device: P ratio = GFLOPs base / GFLOPs pruned (4) If γ corresponding to the feature channel is less than the threshold T, then it is determined as a secondary channel and subjected to pruning to obtain the target teacher model. Specifically, prune the channels with the scaling factor close to zero by deleting all input and output connections and the corresponding weights. Record the weight changes after each pruning to obtain the initial weight sequence; use the initial weight sequence as the training weight of the student model to construct the initial student model, and do not prune the structure with residual connections in the network to ensure that the dimensions of the diameter connection and the residual layer feature map are consistent.
[0047] S3. Adjust the parameters of the initial student model through the knowledge distillation strategy in combination with the target teacher model to obtain the target student model; Specifically, S3 includes the following steps: S31. Initialize the target teacher model and the initial student model so that their channel numbers and feature map sizes are adapted; use the KL divergence to measure the feature differences between the target teacher model and the initial student model on each feature map to obtain the internal hierarchical feature distillation loss value; S32. Calculate the difference between the predicted bounding box of the target teacher model and the predicted bounding box of the initial teacher model based on the shape intersection over union metric to obtain the bounding box distillation loss; Synchronously, measure the difference between the classification prediction of the initial student model and the classification prediction of the target teacher model to determine the classification distillation loss; S33. Calculate the external hierarchical feature distillation loss value through the bounding box distillation loss and the classification distillation loss; perform weighted summation through the internal hierarchical feature distillation loss value, the external hierarchical feature distillation loss value, and their corresponding weight values to obtain the fused feature distillation loss; use the minimum of the fused feature distillation loss as the objective function to solve for the parameter adjustment value, and correct the parameter adjustment value in the initial student model to obtain the target student model.
[0048] In this embodiment, through the knowledge distillation strategy, the knowledge of the target teacher model is transferred to the initial student model. The internal hierarchical feature distillation loss value prompts the initial student model to learn the feature representation of the target teacher model on the feature map. The bounding box distillation loss and the classification distillation loss respectively enable the student model to learn the capabilities of the teacher model in object detection and classification, thereby improving the overall performance of the student model and enabling it to more accurately detect and classify defects in the vibration damping hammer multi-defect detection task. The initial student model has a relatively simple structure and is more suitable for deployment on edge devices. By adjusting the parameters through knowledge distillation, the lightweight of the model is achieved while ensuring the model performance. At the same time, due to the adaptation process of the target teacher model and the initial student model, the finally obtained target student model can better adapt to the computing resources and performance requirements of edge devices, improving the operation efficiency of the model on edge devices. The fused feature distillation loss comprehensively considers the feature differences between the internal and external hierarchies, enabling the target student model to learn the knowledge of the target teacher model from multiple perspectives during the learning process and enhancing the generalization ability of the model. When facing different types and scenarios of vibration damping hammer defects, the target student model can detect and classify more stably and accurately, reducing the risk of overfitting.
[0049] It can be understood that knowledge distillation is a technique in deep learning for transferring the knowledge of a complex model (teacher model) to a simpler model (student model). In this embodiment, the internal hierarchical feature and external logits collaborative knowledge distillation (FLSKD) method is adopted, as Figure 2 shown in the upper half. The blue network above represents the teacher network, and the green network below represents the student network. FLSKD combines the internal hierarchical feature distillation (ILFD) technique and the external logits knowledge distillation (OLKD) technique.
[0050] By coordinating the information of the internal hierarchical features and the final output results during the model training process, the learning efficiency and performance of the student model are improved. The FLSKD loss is shown in Equation (5), where L ILFD is the internal hierarchical feature distillation loss value, L OLKDis the external hierarchical feature distillation loss value. α and β are the weight coefficients of the internal hierarchical feature distillation and the external logits knowledge distillation coefficients respectively. By adjusting these two coefficients, the performance of the internal hierarchical feature distillation and the external logits knowledge distillation can be balanced. The total loss of the model is shown in Equation (6), L total is the total loss during the model training process, L model is the loss during the normal training process of the model; L FLSKD = α·L ILFD + β·L OLKD (5) L total = L model + L FLSKD (6) Model pruning reduces the model complexity and improves the inference speed of the model. However, during the process of model lightweighting, pruning will sacrifice the model's ability to capture features, thereby affecting the detection accuracy. Therefore, compensating for the loss of useful features during the lightweighting process is the key to ensuring the accurate detection of vibration damping hammer defects by the model. The ILFD technology includes two parts: channel-level feature distillation technology and metric learning. By combining channel-level feature distillation and metric learning, the student model can better learn the feature representation of the teacher model, thereby improving the detection performance of the student model.
[0051] First, teacher_f and student_f are the feature lists of the teacher model and the student model respectively. As Figure 3 shown, the initialization of the ILFD module includes an alignment module and a normalization module. The alignment module uses a 1×1 convolutional layer to adjust the number of channels of the student model to ensure the matching of the number of channels between the student model and the teacher model, and at the same time ensure the correspondence of the feature map sizes of the teacher model and the student model. The normalization module is used for the normalization processing of the aligned feature map to make it have the same dynamic range. To achieve more accurate target localization and defect category judgment, this embodiment designs an OLKD module. As Figure 4The OLKD shown includes two parts: bounding box distillation and classification distillation, aiming to transfer the prediction results of the target teacher model to the initial student model so that the initial student model can learn more accurate decision boundaries. First, teacher_p and student_p are the prediction results of the teacher model and the student model respectively, and then the prediction results are split. t_pwh and s_pwh are the widths and heights of the bounding boxes predicted by the teacher model and the student model respectively. t_pxy and s_pxy are the center point coordinates of the bounding boxes predicted by the teacher model and the student model respectively. t_logits and s_logits are the object category logits distributions predicted by the teacher model and the student model respectively. t_pwh, t_pxy, s_pwh, and s_pxy are converted into the actual bounding boxes t_box and s_box that can locate the target on the image through the Decoder (decoder).
[0052] As an alternative embodiment, using the KL divergence to measure the feature differences between the target teacher model and the initial student model on each feature map and then obtaining the internal hierarchical feature distillation loss value includes the following steps: Calculate the Euclidean distance between the feature vector and all other feature vectors in the batch to obtain a distance matrix, and create a mask matrix with the same dimension as the distance matrix according to the class labels to identify samples with the same label; Find the maximum distance among the samples with the same label as the hardest positive sample distance, and the minimum distance among the samples with different labels as the hardest negative sample distance; Calculate the distance loss value between the hardest positive sample and the hardest negative sample through the triplet loss; measure the feature difference with the distance loss to obtain the bounding box distillation loss.
[0053] It is understandable that during the forward propagation process, the KL divergence is used to measure the feature differences between the teacher network and the student network on each feature map, and the channel-level distillation loss is obtained. At the same time, the triplet loss is used to enhance the discriminability of the features of the student model. For each sample in the batch, the Euclidean distances between the feature vector of this sample and the feature vectors of all other samples in the batch are calculated to form a distance matrix. Then, a mask matrix with the same dimension as the distance matrix is created according to the class labels to identify which samples have the same label. For each sample, the maximum distance among the samples with the same label as it is found as the hardest positive sample distance, and the minimum distance among the samples with different labels from it is found as the hardest negative sample distance. The distance loss between the hardest positive sample and the hardest negative sample is calculated through the triplet loss, aiming to make the distance of the hardest negative sample greater than the distance of the hardest positive sample plus a margin m, so as to ensure that the samples of the same category are closer in the feature space, while the samples of different categories are farther apart. It helps the student model learn a feature space that is compact within each category and separated between categories during the distillation process. The ILFD module compensates for the decline in the feature capture ability of the pruned model. However, due to the complex background environment where the vibration dampers are located and the small scale of the vibration damper defects in the real environment, it is difficult for the model to accurately determine the decision boundary for small targets.
[0054] As an alternative embodiment, determining the classification distillation loss by measuring the difference between the classification prediction of the initial student model and the classification prediction of the target teacher model includes the following steps: The classification logits of the target teacher model and the initial student model are normalized through sigmoid, and the binary cross-entropy loss between the classification prediction of the initial student model and the classification prediction of the teacher model is calculated to obtain the classification prediction difference to determine the classification distillation loss.
[0055] It can be understood that, as shown in formula (10), this function measures the difference between the predicted bounding boxes of the student model and those of the teacher model. Specifically, the bounding box distillation loss is calculated based on the Shape-IoU (Shape Intersection over Union) metric, which quantifies the spatial overlap between the predicted bounding boxes of the student model and those of the teacher model. Different from IoU, Shape-IoU takes into account the shape information of the object and compares the shape similarity of two objects. While IoU only considers the ratio between the intersection and union of two bounding boxes and does not consider shape information. Therefore, Shape-IoU can more accurately measure the similarity between objects, and by calculating the loss by focusing on the bounding box shape, the bounding box regression of the student model is more accurate. Among them, Shape-IoUi represents the Shape-IoU between the predicted bounding box of the i-th student model and the predicted bounding box of the teacher model, and n represents the number of all sample points in each image. By minimizing the values of all inverse Shape-IoU, it aims to make the bounding box prediction of the student model closer to that of the teacher model, thereby improving the accuracy of object localization of the student model; For classification distillation, in order to measure the difference between the classification predictions of the student model and those of the teacher model, the classification logits map is converted into multiple binary classification maps during the distillation process. The classification logits of the target teacher model and the initial student model are normalized through sigmoid, and then the binary cross-entropy loss between the classification predictions of the initial student model and the classification predictions of the target teacher model is calculated, making the classification predictions of the student model as close as possible to those of the teacher model.
[0056] The expression of the classification distillation loss function is as shown in formula (11). Among them represents the result of the sigmoid of the student model classification s_logits, represents the result of the sigmoid of the teacher model classification t_logits. n is the number of sample points, K is the number of categories, LBCE is the binary cross-entropy loss, and Lcls is the distillation loss. By minimizing the classification distillation loss, the student model can better learn the distribution characteristics of the target categories, thereby improving the performance of object recognition; The final OLKD loss is composed of the bounding box distillation loss and the classification distillation loss, as shown in formula (12); L OLKD =L box +L cls (12).
[0057] S4. The target student model adapts to the target scenario sequence to obtain an edge configuration path, and implants the target student model into the corresponding edge device according to the edge configuration path.
[0058] Specifically, it includes the following steps: S41. Set a scenario adaptation threshold, filter out the target scenario sequences greater than the scenario adaptation threshold, and calculate the rated working computing power values required for each fault scenario; determine the edge devices greater than the rated working computing power values as preselected edge devices; S42. Determine the model computing power value of the target student model according to the distillation loss, use the difference between the model computing power value and the rated working computing power value as the numerator, and use the difference between the actual maximum computing power value of the rated working computing power value and the rated working computing power value as the denominator; obtain the computing power resource utilization rate, and use the preselected edge device with the maximum computing power resource utilization rate as the target edge device; S43. Construct a model receiving channel for the target edge device through an asymmetric encryption algorithm to obtain an edge configuration path, and the target student model is sent to the corresponding target edge device through the edge configuration path.
[0059] In this embodiment, by setting a scenario adaptation threshold to filter the target scenario sequences, calculating the rated working computing power value and the model computing power value, and selecting the edge device with the maximum computing power resource utilization rate as the target edge device, the computing resources of the edge device can be fully utilized, avoiding over-allocation or idleness of resources, and improving the resource utilization efficiency of the entire power transmission network operation and maintenance system. For example, assume that in the above scenario, the edge device with the most suitable 800 GFLOPS computing power is selected, so that the resources are more reasonably allocated. Using an asymmetric encryption algorithm to construct a model receiving channel ensures the security and integrity of the target student model during transmission. Even if the data is intercepted during transmission, the attacker cannot decrypt the data without the private key, preventing the leakage and tampering of the model data, ensuring that the model can be accurately deployed to the target edge device, and improving the reliability and stability of the system. Adapting the target student model to the target scenario sequence can make the model better adapt to the requirements of different fault scenarios. Deploying the model to the most suitable edge device can give full play to the performance advantages of the edge device, improve the accuracy and efficiency of anti-vibration hammer multi-defect detection, timely discover and handle problems in the power transmission line, and ensure the stable operation of the power transmission line.
[0060] The specific means of implanting the target student model into the edge device are as follows: For example, it is the process of converting the ONNX model into a TensorRT engine for model optimization. First, use the C++ API provided by TensorRT to convert the ONNX model into a format recognizable by TensorRT, achieving the purpose of loading the ONNX model into TensorRT. Next, configure the TensorRT engine to optimize the performance of the engine by setting parameters such as network structure, inference batch size, and inference precision. After the configuration is completed, compile the TensorRT engine into an executable file. Finally, at runtime, use the compiled TensorRT engine for inference operations. Through this process, the optimization capabilities of TensorRT can be fully utilized to achieve efficient deep learning inference on Jetson Xavier NX.
[0061] It can be understood that assuming in a large power transmission network management area, the target scene sequence and the target student model have been obtained. Set the scene fitness threshold to 0.7, and screen out the sequences with a scene fitness greater than 0.7 among the numerous target scene sequences. For example, there are 10 different fault scenarios, and each scenario corresponds to a set of edge devices and related parameters. After screening, 6 scenarios have a scene fitness greater than 0.7. For these 6 fault scenarios, calculate the rated working computing power value required for each scenario. Taking one of the fault scenarios as an example, this scenario is to detect multiple cracks and component looseness defects of the vibration damping hammer. According to factors such as the complexity of the defects, image resolution, and model processing flow, after multiple tests and analyses, it is determined that the rated working computing power value required for this fault scenario is 500 GFLOPS (giga floating-point operations per second).
[0062] Furthermore, check the actual computing power of the edge devices in the area corresponding to each fault scenario. Assume there are 5 edge devices near this fault scenario, and their actual maximum computing power values are 400 GFLOPS, 600 GFLOPS, 700 GFLOPS, 550 GFLOPS, and 800 GFLOPS respectively. Determine the edge devices with an actual maximum computing power value greater than 500 GFLOPS, that is, the 4 edge devices with computing powers of 600 GFLOPS, 700 GFLOPS, 550 GFLOPS, and 800 GFLOPS as the preselected edge devices.
[0063] Determine the model computing power value according to the distillation loss of the target student model during the training process. Suppose that after statistics and analysis, the model computing power value of the target student model when processing this fault scenario is 450 GFLOPS. For each preselected edge device, calculate its computing power resource utilization rate. Take the edge device with a computing power of 600 GFLOPS as an example. According to the formula: computing power resource utilization rate = (model computing power value - rated working computing power value) / (actual maximum computing power value - rated working computing power value), that is, (450 - 500) / (600 - 500) = -0.5 (this is just an example to illustrate the calculation method. In actual situations, there may be more reasonable calculations and processing because the appearance of negative numbers may indicate some special situations of computing power requirements and allocations, and cannot be used as a limitation of this embodiment); for the edge device with a computing power of 700 GFLOPS, the computing power resource utilization rate = (450 - 500) / (700 - 500) = -0.25; for the edge device with a computing power of 550 GFLOPS, the computing power resource utilization rate = (450 - 500) / (550 - 500) = -1; for the edge device with a computing power of 800 GFLOPS, the computing power resource utilization rate = (450 - 500) / (800 - 500) ≈ -0.17. After comparison, select the edge device with the largest (within a reasonable range) computing power resource utilization rate as the target edge device. Suppose that after comprehensively considering various factors, the edge device with a computing power of 800 GFLOPS is selected as the target edge device.
[0064] Furthermore, in order to ensure the safe and accurate transmission of the target student model to the target edge device, an asymmetric encryption algorithm is used to construct a model receiving channel. For example, using the RSA asymmetric encryption algorithm, the edge device generates a pair of public and private keys and sends the public key to the model deployment end. The model deployment end uses this public key to encrypt the target student model to generate encrypted model data. Then, the encrypted model data is sent to the target edge device through the network. The target edge device uses its own private key to decrypt the encrypted data to obtain the target student model. In this way, the construction of the edge configuration path is completed, and the target student model is sent to the corresponding target edge device through this path.
[0065] The above - described specific implementation manner is a preferred implementation manner of the edge configuration method of the lightweight deployment model based on the multi - defect detection of vibration dampers of the present invention, and does not limit the specific scope of the present invention. The scope of the present invention includes but is not limited to this specific implementation manner. All equivalent changes made according to the shape and structure of the present invention are within the protection scope of the present invention.
Claims
1. A lightweight deployment model edge configuration method based on multi-defect detection of vibration dampers, characterized in that: It includes the following steps: S1. Perform scene similarity processing on the anti-vibration hammer multi-defect dataset and the edge device attribute dataset to obtain the target scene sequence; S2. Sequentially perform sparsification training on the initial teacher model through the target scene sequence to obtain the target teacher model and the corresponding initial weight sequence; construct the initial student model based on the initial weight sequence; S3. Adjust the parameters of the initial student model through the knowledge distillation strategy in combination with the target teacher model to obtain the target student model; S4. The target student model adapts to the target scene sequence to obtain the edge configuration path, and implants the target student model into the corresponding edge device according to the edge configuration path.
2. The lightweight deployment model edge configuration method based on anti-vibration hammer multi-defect detection according to claim 1, characterized in that: The step of performing scene similarity processing on the anti-vibration hammer multi-defect dataset and the edge device attribute dataset to obtain the target scene sequence includes the following steps: S11. Obtain the anti-vibration hammer defect images in the power transmission management domain to construct the anti-vibration hammer multi-defect dataset, and obtain all edge device information to construct the edge device attribute dataset; S12. Cluster the elements in the anti-vibration hammer multi-defect dataset according to defect similarity to obtain the defect sequence; Synchronously, cluster the elements in the edge device attribute dataset according to resource similarity to obtain the resource sequence; S13. Extract the edge device information corresponding to the defect scene, and integrate the defect sequence and the resource sequence to obtain the target scene sequence.
3. The lightweight deployment model edge configuration method for multi-defect detection based on vibration dampers according to claim 2, characterized in that: The S12 includes the following steps: Extract the color features, texture features, and shape features of the anti-vibration hammer defect images, calculate the cosine similarity between all images in the anti-vibration hammer multi-defect dataset through the cosine similarity algorithm, and divide the images with cosine similarity higher than the set threshold into the same category to obtain the defect sequence.
4. The lightweight deployment model edge configuration method for multi-defect detection based on vibration dampers according to claim 2, characterized in that: The S12 further includes the following steps: Extract the computing power features, storage capacity features, and network bandwidth features of the edge devices, calculate the Euclidean distance between all edge devices in the edge device attribute dataset through the Euclidean distance algorithm, and divide the edge devices with Euclidean distance less than the set threshold into the same category to obtain the resource sequence.
5. The lightweight deployment model edge configuration method based on multi-defect detection of vibration dampers according to claim 2, characterized in that: The step of extracting the edge device information corresponding to the defect scene and integrating the defect sequence and the resource sequence to obtain the target scene sequence includes the following steps: S131. Search for all edge devices in the area centered on the defect scene with a search radius of r; S132. Calculate the scene adaptation degree according to the connection degree, communication distance, and resource matching degree between the edge device and the defect scene; S133. Sort the edge devices according to the scene adaptation degree to determine the target scene sequence corresponding to each defect scene.
6. The lightweight deployment model edge configuration method based on vibration damper multi-defect detection according to claim 1, characterized in that: Sequentially perform sparsification training on the initial teacher model through the target scene sequence to obtain the target teacher model and the corresponding initial weight sequence; construct the initial student model based on the initial weight sequence; includes the following steps: S21. Train the initial teacher model through the target scene sequence to obtain the batch processing dataset, perform normalization processing on the batch processing dataset, and correct the scaling factor through the preset scaling factor and bias parameters of the channel to obtain the target scaling factor; S22. Sparsely represent the set Γ of target scaling factors corresponding to each channel in the convolutional layer by using L1 regularization; Connect the sparsely trained target scaling factors with the corresponding channels in the convolutional layer feature map, and sort them in absolute value to obtain a joint sequence; S23. Determine the pruning threshold according to the required computing power value of the basic model and the actual maximum computing power value of the edge device; If the target scaling factor corresponding to the feature channel is less than the pruning threshold, it is determined as a secondary channel and the joint sequence is trimmed to obtain the target teacher model; S24. Record the weight changes after each pruning to obtain the initial weight sequence; use the initial weight sequence as the training weight of the student model to construct the initial student model.
7. The lightweight deployment model edge configuration method based on vibration damper multi-defect detection according to claim 1, characterized in that: The initial student model is adjusted in parameters by using the knowledge distillation strategy in combination with the target teacher model to obtain the target student model; it includes the following steps: S31. Initialize the target teacher model and the initial student model so that their channel numbers and feature map sizes are adapted; use KL divergence to measure the feature differences between the target teacher model and the initial student model on each feature map to obtain the internal hierarchical feature distillation loss value; S32. Calculate the difference between the predicted bounding box of the target teacher model and the predicted bounding box of the initial teacher model based on the shape intersection over union metric to obtain the bounding box distillation loss; Synchronously, measure the difference between the classification prediction of the initial student model and the classification prediction of the target teacher model to determine the classification distillation loss; S33. Calculate the external hierarchical feature distillation loss value through the bounding box distillation loss and the classification distillation loss; perform weighted summation through the internal hierarchical feature distillation loss value, the external hierarchical feature distillation loss value and their corresponding weight values to obtain the fused feature distillation loss; solve with the minimum of the fused feature distillation loss as the objective function to obtain the parameter adjustment value, and correct the parameter adjustment value in the initial student model to obtain the target student model.
8. The lightweight deployment model edge configuration method for multi-defect detection based on vibration dampers according to claim 7, characterized in that: The step of using KL divergence to measure the feature differences between the target teacher model and the initial student model on each feature map to obtain the internal hierarchical feature distillation loss value includes the following steps: Calculate the Euclidean distance between the feature vector and all other feature vectors in the batch to obtain a distance matrix, and create a mask matrix with the same dimension as the distance matrix according to the class labels to identify the samples with the same label; Find the maximum distance among the samples with the same label as the hardest positive sample distance, and the minimum distance among the samples with different labels as the hardest negative sample distance; Calculate the distance loss value between the hardest positive sample and the hardest negative sample through the triplet loss; measure the feature differences with the distance loss to obtain the bounding box distillation loss.
9. The edge configuration method for lightweight deployment model based on multi-defect detection of vibration dampers according to claim 7, wherein: The step of measuring the difference between the classification prediction of the initial student model and the classification prediction of the target teacher model to determine the classification distillation loss includes the following steps: The classification logits of the target teacher model and the initial student model are normalized by sigmoid, and the binary cross-entropy loss between the classification prediction of the initial student model and the classification prediction of the teacher model is calculated to obtain the classification prediction difference to determine the classification distillation loss.
10. The lightweight deployment model edge configuration method based on multi-defect detection of vibration dampers according to claim 1 or 5, characterized in that: The target student model adapts to the target scenario sequence to obtain an edge configuration path, and implants the target student model into the corresponding edge device according to the edge configuration path, including the following steps: S41. Set a scenario adaptation threshold, screen out the target scenario sequences greater than the scenario adaptation threshold, and calculate the required rated working computing power value for each fault scenario; determine the edge devices greater than the rated working computing power value as preselected edge devices; S42. Determine the model computing power value of the target student model according to the distillation loss, use the difference between the model computing power value and the rated working computing power value as the numerator, and use the difference between the actual maximum computing power value of the rated working computing power value and the rated working computing power value as the denominator; Obtain the computing power resource utilization rate, and use the preselected edge device with the maximum computing power resource utilization rate as the target edge device; S43. Construct a model receiving channel for the target edge device through an asymmetric encryption algorithm to obtain an edge configuration path, and the target student model is sent to the corresponding target edge device through the edge configuration path.
Citation Information
Patent Citations
Lightweight power transmission line insulator defect detection method based on knowledge distillation
CN119228767A
Cited By
Light-weight man-machine interaction model system for industrial scene application
CN121213892A