Large-model cross-end transfer compression method and device applied to intelligent electric energy meter

The method addresses heterogeneous device challenges in smart meter clusters by using device-aware compression and collaborative transmission to enhance resource utilization and reduce misreporting, ensuring efficient deployment and real-time anomaly detection.

CN120321130AActive Publication Date: 2025-07-15CSG SMART SCI&TECH CO LTD +1

Patent Information

Application Number
CN202510788677.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-15
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

The existing technology cannot effectively solve the problems of cross-end migration complexity of model, low resource utilization, poor group collaboration efficiency and large feature alignment errors caused by equipment heterogeneity in smart power meters clusters.

Method used

Using the methods of device feature-aware compression, dynamic adapter injection, and collaborative differential transmission, the feature fingerprint is generated through quantitative modeling, the model structure is dynamically adjusted, and the closed-loop process of terminal feature analysis-hierarchical dynamic compression-cluster collaborative distribution-online fine-tuning and calibration is built to achieve accurate adaptation between the model and the terminal environment.

Benefits of technology

It improves resource utilization, reduces false alarm rate, improves model compression efficiency and accuracy, reduces redundant communication traffic, and adapts to real-time calibration requirements in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321130A_ABST
    Figure CN120321130A_ABST
Patent Text Reader

Abstract

The invention discloses a large-model cross-end transfer compression method applied to an intelligent electric energy meter, and the method comprises the steps: S1, carrying out the quantitative modeling of the hardware characteristics of terminal equipment, and generating a feature fingerprint for guiding the compression of a model based on the quantitative data; s2, a compression strategy of each layer of a model structure after quantitative modeling is dynamically adjusted based on fingerprints, and the resource utilization rate is maximized on the premise that precision is guaranteed; s3, realizing efficient distribution of cluster-level model update by constructing a transmission network of equipment topology perception; and S4, a lightweight adaptation module is injected into the terminal, the feature difference between the devices is rapidly eliminated, and accurate adaptation of the model and the terminal environment is completed. According to the method, through three core technologies of equipment feature perception compression, dynamic adapter injection and collaborative differential transmission, a full-process closed loop of terminal feature analysis-layered dynamic compression-cluster collaborative distribution-online fine tuning calibration is constructed, and a bottleneck breakthrough of cross-end migration of a heterogeneous terminal cluster model can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart grid edge computing, and particularly relates to a large model cross-terminal transfer compression method and device applied to smart electricity meters, which are applicable to collaborative anomaly detection and dynamic model optimization of smart electricity meter groups. Background Art

[0002] Anomaly detection and model optimization of smart electricity meter clusters are one of the core technologies of power system intelligence. Traditional methods mainly rely on cloud centralized processing to analyze meter data and identify faults through a unified deployed deep learning model. However, with the heterogeneity of terminal devices (such as differences in processor architectures, wide range of memory capacities, and varying sensor accuracies) and the expansion of the deployment scale, the existing technologies face significant challenges. On the one hand, cloud models are difficult to directly adapt to low-power terminals (such as Cortex-M7, RISC-V, etc.), resulting in low model compression, transmission, and operation efficiency; on the other hand, hardware differences between terminal devices (such as computing power ranging from 0.1 TOPS to 2 TOPS and memory ranging from 256 KB to 2 GB) further exacerbate the complexity of model cross-terminal migration. Therefore, there is an urgent need for a cross-terminal migration compression method for power large models facing heterogeneous terminals to balance model performance and terminal resource limitations.

[0003] To address the above problems, existing technologies usually adopt static hierarchical compression technology and collaborative distillation technology to solve them. However, although the static hierarchical compression technology can adjust the model through preset parameters (such as fixed pruning rate, unified quantization bit width), it lacks the ability to perceive the dynamic characteristics of terminals (computing power, memory, sensor accuracy); at the same time, although the collaborative distillation technology can rely on edge gateways for subgraph generation and parameter synchronization, it does not solve the problem of model performance degradation caused by inconsistent feature distributions between devices (such as ADC accuracy differences); in addition, there are also conventional transmission optimization operations, which although attempt to reduce redundant traffic through model distribution strategies, do not combine device topology structures to design multi-hop routing and difference aggregation mechanisms.

[0004] Therefore, the core limitations of these methods are: unable to dynamically perceive terminal heterogeneity and adaptively adjust compression strategies, resulting in low resource utilization, poor group collaboration efficiency, and large feature alignment errors.

[0005] In addition, in the prior art, static hierarchical compression technology and collaborative distillation technology are two typical solutions. For example: (1) Static hierarchical compression technology generally realizes model lightweighting through global pruning (such as a fixed pruning rate of 50%), unified quantization (such as 8-bit fixed-point quantization), and unicast transmission (the model is independently sent down through the MQTT protocol). However, its defect is that: since the pruning rate of high-computing-power devices (above 1 TOPS) is limited by the preset ratio, it is easy to cause 60% of the computing resources to be idle. In addition, technical problems such as memory mismatch and redundant transmission are likely to occur; (2) Collaborative distillation technology generally realizes group optimization through edge gateway generation of lightweight subgraphs, gradient aggregation, and version management. However, its defect is that: problems such as feature drift, excessive energy consumption, and convergence lag are likely to occur. Among them, during the convergence lag process, the group model needs more than 20 rounds of iteration to be stable, which is not applicable to delay-sensitive scenarios.

[0006] Therefore, when facing device heterogeneity, resource dynamic changes, and group collaboration requirements, the above technical solutions all have problems such as low efficiency and excessive energy consumption, and it is difficult to support the real-time anomaly detection requirements of large-scale intelligent electricity meter clusters.

[0007] For this reason, this application specifically proposes a large model cross-terminal transfer compression method applied to intelligent electricity meters to solve the above technical problems. Summary of the Invention

[0008] The main object of the present invention is to provide a large model cross-terminal transfer compression method applied to intelligent electricity meters. Through three core technologies of device feature perception compression, dynamic adapter injection, and collaborative differential transmission, a full-process closed loop of "terminal feature analysis - hierarchical dynamic compression - cluster collaborative distribution - online fine-tuning and calibration" is constructed to solve the technical problems proposed in the background technology.

[0009] The present invention adopts the following technical solutions to solve the above technical problems: A large model cross-terminal transfer compression method applied to intelligent electricity meters includes: S1. Quantitatively model the hardware characteristics of the terminal device, and at the same time, based on the quantization data, generate a feature fingerprint for guiding model compression; S2. Dynamically adjust the compression strategy of each layer of the model structure after quantization modeling based on the fingerprint, including: using the feature fingerprint to parse and identify the computationally intensive layer, memory-sensitive layer, and precision-sensitive layer of the model structure; And sequentially execute operations: dynamically adjust the pruning rate of the computationally intensive layer according to the device computing power index, perform differential and quantization operations on the memory-sensitive layer, and perform hybrid quantization on the precision-sensitive layer to retain the FP16 precision of the key channels, ensuring that the harmonic feature extraction error is less than 0.5%; S3. Realize the efficient distribution of cluster-level model updates by constructing a transmission network that perceives the device topology; S4. Inject a lightweight adaptation module into the terminal to quickly eliminate the feature differences between devices and complete the precise adaptation of the model to the terminal environment.

[0010] Preferably, during the quantization modeling process in S1, the cloud analyzes device parameters including the computing power index, memory adaptability, and precision correction factor to construct a multi-dimensional feature vector, complete the modeling quantization operation, and use the hierarchical clustering algorithm to divide terminals with similar features into isomorphic groups, where: The formula for the computing power index is:

[0011] Where, is the main frequency of the processor of the terminal device, is the number of cores of the terminal device, is the number of instructions per cycle of the terminal device; The formula for evaluating the memory adaptability is:

[0012] Where, is the available memory of the terminal, is the number of parameters of the benchmark model, is the total memory value of the terminal, and indicates that the memory is sufficient, indicates that the memory is insufficient; The formula for the precision correction factor is:

[0013] Where, is the RMS noise of the ADC, is the ADC resolution. Preferably, the specific operation process of the operation in S2 includes: Dynamically adjust the pruning rate of the compute-intensive layer according to the device computing power index, where the pruning rate formula is:

[0014] Where, represents the value of the device with the highest computing power in the cluster, is the value of this computing power device, and there is a constraint condition: ; Perform differential and quantization operations on the memory-sensitive layer. At this time, the quantization bit width allocation amount of the memory-sensitive layer is:

[0015] Where, is the maximum memory adaptation value in the cluster, is the memory adaptation value of the computing power device, and there is a lower bound constraint: , is expressed as the floor symbol, used to represent the largest integer not greater than ; For the mixed quantization precision sensitive layer, there are:

[0016] Among them, is the quantization step, used to retain the FP16 precision of the key channels, and force the mixed quantization to be enabled when .

[0017] Preferably, the specific operation process of performing model update distribution through the transmission network in S3 includes: S31. Elect a terminal with edge computing capabilities as the proxy node in each device group; S32. The cloud compares the old and new model versions, extracts the parameter change amounts of each layer, and only retains the core parameters with a change amplitude > 2%; S33. According to the geographical distribution of the device groups and the communication link quality, construct a multi-hop transmission path, where the proxy node preferentially receives the complete update packet, and then broadcasts it to the group members through the local area network. When transmitting across groups, select the optimal path with a relay hop count ≤ 3 to avoid network congestion; S34. After the terminal receives the update packet, use the hash check and rollback mechanism until the error rate < 0.01% to ensure the transmission integrity.

[0018] Preferably, the specific transmission process of the transmission network for performing efficient data transmission in S3 includes: L1. Generate a difference parameter matrix. The matrix generation formula between nodes and is:

[0019] Among them, and respectively represent the old and new model parameter matrices between nodes and ; After that, only retain the parameters of ; L2. Calculate the transmission priority. The calculation formula is:

[0020] Among them, represents the transmission priority, represents the device group The number of hops to the cloud, represented as the size of the device group ; represented as the group Frobenius norm of the parameter difference matrix; L3. The proxy node performs difference aggregation, and there is:

[0021]

[0022] Among them, represented as the aggregated difference parameter, represented as the top with the largest magnitude among the selected single-device difference parameters, represented as the device device difference parameter during transmission, represented as the set of device groups, represented as the floor symbol, used to represent the largest integer not greater than .

[0023] Preferably, the specific operation process of the lightweight adaptation module in S4 to eliminate feature differences includes: S41. Generate an adapter weight matrix based on the device feature fingerprint, and the formula is:

[0024] Among them, is the adapter weight matrix after weight initialization, represents a two-layer fully connected network with the input being the device feature vector ; S42. After the terminal loads the compressed model, freeze the parameters of the backbone network and iteratively adjust the adapter; During the adjustment process, adjust based on the loss function of the adapter, and there is:

[0025] Among them, is the loss function of the adapter, is the balance coefficient, and linearly decays from 1.0 to 0.7, is the cross-entropy loss function, is the adapter weight matrix after the previous round of adjustment, initially is 0, is the square of the L2 norm; During the adjustment process, use local historical data for forward inference and loss calculation, and prevent overfitting by restricting the gradient update amplitude; S43. Dynamically correct the feature activation threshold according to real-time environmental parameters, where the threshold adjustment rule is:

[0026] where is the basic feature activation threshold, is the updated and adjusted feature activation threshold, is the temperature sensitivity coefficient, is the nominal temperature, is the implementation temperature.

[0027] Preferably, it further includes S5: Regularly report operating metrics including model inference latency, memory occupancy rate, and false alarm rate through the terminal, analyze cluster-level data through the cloud, identify devices with abnormal performance, and dynamically adjust the next-round compression parameters according to the feedback data. The specific operation process of dynamically adjusting the next-round compression parameters according to the feedback data includes: S51. Construct an iterative formula for the compression strategy, as follows:

[0028] where is the weight of the operating metric; S52. Optimize the transmission path, using the transmission path evaluation as the optimization criterion, as follows;

[0029] where is the optimized transmission path evaluation, is the bandwidth path of the th hop of the path, is the maximum allowable bandwidth path of the path, is the end-to-end delay of the th hop of the path, is the maximum allowable end-to-end delay.

[0030] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.

[0031] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.

[0032] As can be seen from the above technical solutions, the present invention provides a large model cross-terminal transfer compression method applied to intelligent electricity meters. Compared with the prior art, the present invention has the following advantages: 1. The present invention can dynamically adapt to terminal heterogeneity by setting up a mechanism for quantifying and modeling terminal hardware parameters and clustering and grouping in the device feature fingerprint coding, thereby improving resource utilization rate and reducing false alarm rate, and providing a differential model compression strategy for heterogeneous device clusters.

[0033] 2. By setting up computing power-oriented compression, memory-oriented compression, and precision fidelity compression strategies in hierarchical dynamic compression, the present invention can maximize the utilization rate of terminal computing resources and memory resources, thereby improving model compression efficiency and the utilization rate of high-computing-power device resources to ensure the accuracy requirements of key tasks (such as harmonic detection).

[0034] 3. By setting up mechanisms for proxy node election, difference parameter extraction, and topology routing optimization in collaborative differential transmission, the present invention can reduce redundant communication traffic and compress the time-consuming of cluster updates, facilitating the efficient distribution and deployment of large-scale cluster models.

[0035] 4. By setting up lightweight adapter injection and dynamic threshold adjustment strategies in online fine-tuning and calibration, the present invention can quickly eliminate feature differences between devices, suppress the feature alignment error within 5%, and greatly reduce the group false alarm rate, facilitating the adaptation to real-time calibration requirements in complex environments.

[0036] 5. By setting up a mechanism for collecting performance indicators and iterating compression strategies in closed-loop feedback optimization, the present invention can dynamically optimize model compression parameters to improve the quantization accuracy of high false alarm rate device groups and achieve a dynamic balance between model performance and terminal resource constraints.

[0037] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The schematic diagrams of the drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is the overall process schematic diagram of the present invention; Figure 2 is the schematic diagram of the operation detail process of the present invention; Figure 3 is the schematic diagram of the data processing process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0040] In the embodiment, refer in detail to Figures 1 to 3 .

[0041] As Figures 1 to 3 shown. The large model cross-terminal transfer compression method applied to the intelligent electricity meter proposed in the embodiment of the present invention constructs a full-process closed loop of "terminal feature analysis - hierarchical dynamic compression - cluster collaborative distribution - online fine-tuning and calibration" through three core technologies: device feature perception compression, dynamic adapter injection, and collaborative differential transmission, including the following steps: S1. Quantitatively model the hardware characteristics of the terminal device, and at the same time, based on the quantitative data, generate a feature fingerprint for guiding model compression to construct a device feature fingerprint code.

[0042] During the terminal feature collection process, each electricity meter actively reports hardware configuration parameters, including key attributes such as processor architecture (ARM / RISC-V), memory capacity (256KB - 2GB), ADC sampling accuracy (12 - 16bit), communication module type (4G / HPLC), etc.

[0043] During the quantitative modeling process, the cloud analyzes device parameters including computing power index, memory adaptability, and precision correction factor to construct a multi-dimensional feature vector, complete the modeling quantization operation, and uses the hierarchical clustering algorithm to divide terminals with similar features into isomorphic groups, where: The computing power index is the theoretical peak computing power (TOPS) calculated based on the processor main frequency and instruction set, and its calculation formula is:

[0044] Among them, is the processor main frequency of the terminal device (unit: GHz), is the number of cores of the terminal device, is the number of instructions per cycle of the terminal device (take 1.5 for ARM Cortex-M7 and 1.2 for RISC-V); The memory adaptability is used to evaluate the deployment feasibility according to the ratio of the model parameter quantity to the memory capacity, and its evaluation calculation formula is:

[0045] Among them, is the available memory of the terminal (unit: KB), is the number of parameters of the benchmark model (unit: KB), is the total memory value of the terminal, and indicates sufficient memory, indicates insufficient memory; Precision correction factor, used to calibrate the signal acquisition error according to the ADC resolution and noise factor, and its calculation formula is:

[0046] where, is the RMS noise of the ADC (unit: LSB), is the ADC resolution (e.g., 12-bit corresponds to 4096 LSB).

[0047] In addition, it should be noted that in the process of device clustering and grouping, the hierarchical clustering algorithm can also be used to divide terminals with similar features into homogeneous groups (such as high computing power group, low memory group, etc.), providing a basis for subsequent differential compression.

[0048] At this time, by setting the quantization modeling and clustering and grouping mechanism of the terminal hardware parameters in the device feature fingerprint coding, the terminal heterogeneity can be dynamically adapted, thereby improving the resource utilization rate and reducing the false alarm rate, and providing a differential model compression strategy for the heterogeneous device cluster.

[0049] S2. Dynamically adjust the compression strategy of each layer of the model structure after quantization modeling based on the fingerprint, and maximize the resource utilization rate on the premise of ensuring the accuracy.

[0050] The specific operation process of the dynamic adjustment of the compression strategy includes: S21. The feature fingerprint is used as the sensitive layer of the model to parse and identify the computationally intensive layers (such as the LSTM time series prediction layer), memory sensitive layers (such as the fully connected classification layer), and precision sensitive layers (such as the harmonic detection head) of the model structure after quantization modeling; S22. Sequentially perform the operations of dynamically adjusting the pruning rate of the computationally intensive layer according to the device computing power index, the differential and quantization operations on the memory sensitive layer, and the hybrid quantization operation on the precision sensitive layer. The hybrid quantization of the precision sensitive layer is used to retain the FP16 precision of the key channels to ensure that the harmonic feature extraction error < 0.5%. The specific operation process includes: (1) Computing power-oriented compression: Dynamically adjust the pruning rate of the computationally intensive layer according to the device computing power index (the pruning rate of high computing power devices ≤ 40%, and that of low computing power devices ≤ 70%). The pruning rate calculation formula is:

[0051] where, represents the Value (unit: TOPS), is the value of the computing power device, and there are constraint conditions: ; (2) Memory-oriented compression: Differential and quantization operations on memory-sensitive layers (8-bit quantization for high-memory devices, 4-bit hybrid quantization for low-memory devices). At this time, the quantization bit-width allocation for memory-sensitive layers is:

[0052] Among them, is the maximum memory adaptation value in the cluster, is the memory adaptation value of the computing power device, and there is a lower limit constraint condition: (4-bit quantization); (3) Precision-fidelity compression: Hybrid quantization of precision-sensitive layers (retaining the FP16 precision of key channels in precision-sensitive layers to ensure that the harmonic feature extraction error < 0.5%). There is:

[0053] Among them, is the quantization step size, which is used to retain the FP16 precision of key channels and forcibly enables hybrid quantization when ; S23. Insert a micro learnable adapter (parameter quantity < 500) at the model input / output end to compensate for sensor differences between devices.

[0054] In a specific embodiment, a hierarchical compression mechanism driven by device feature fingerprints is constructed. According to the terminal computing power (0.1 - 2 TOPS), memory margin (30 - 512 KB), and sensor accuracy (±0.5 - 2 LSB), the pruning rate (20 - 70%) and quantization bit-width (4 - 16 bit) are dynamically adjusted, so that the computing utilization rate of high-end devices is increased to more than 80%, and the model loading success rate of low-end devices > 98%. Therefore, by setting the computing power-oriented compression, memory-oriented compression, and precision-fidelity compression strategies in hierarchical dynamic compression at this time, the utilization rate of terminal computing resources and memory resources can be maximized. For example, in a heterogeneous device cluster, the average model compression rate reaches 85%, and the resource utilization rate of high-computing power devices is increased to 82%, thereby improving the model compression efficiency and the resource utilization rate of high-computing power devices to ensure the accuracy requirements of key tasks (such as harmonic detection).

[0055] S3. By constructing a device topology-aware transmission network, efficient distribution of cluster-level model updates is achieved.

[0056] Among them, the specific operation process of performing model update distribution through the transmission network includes: S31. Elect a terminal with edge computing capabilities (such as a device with memory ≥ 512KB) as the proxy node in each device group; S32. The cloud compares the old and new model versions, extracts the parameter change amounts of each layer, and only retains the core parameters with a change amplitude > 2%; S33. Construct a multi-hop transmission path according to the geographical distribution and communication link quality of the device groups. The proxy node preferentially receives the complete update packet and then broadcasts it to the group members through the local area network. When transmitting across groups, select the optimal path with a relay hop count ≤ 3 to avoid network congestion; S34. After the terminal receives the update packet, use the hash check and rollback mechanism until the error rate < 0.01% to ensure transmission integrity.

[0057] At this time, the specific transmission process for the transmission network to perform efficient data transmission includes: L1. Generate a difference parameter matrix. The matrix generation formula between node and is:

[0058] Among them, and respectively represent the old and new model parameter matrices between nodes and ; After that, only retain the parameters of ; L2. Calculate the transmission priority. The calculation formula is:

[0059] Among them, represents the transmission priority, represents the number of hops from device group to the cloud (such as when the gateway is the relay ), represents the size of device group , represents the Frobenius norm of the parameter difference matrix of group ; L3. The proxy node performs difference aggregation, and there is:

[0060]

[0061] Among them, represents the aggregated difference parameter, represents the top with the largest amplitude among the selected single-device difference parameters, Denoted as a device Device difference parameters during transmission, Denoted as a set of device groups.

[0062] In a specific embodiment, a topology-aware cooperative transmission protocol is designed. Through proxy node difference aggregation and multi-hop routing optimization, the time-consuming for model update of the cluster is compressed from 47 minutes to 18 minutes, and the communication traffic is reduced to 38% of the traditional scheme. Therefore, by setting up a proxy node election, difference parameter extraction, and topology routing optimization mechanism in cooperative differential transmission, redundant communication traffic can be reduced and the cluster update time-consuming can be compressed, facilitating the efficient distribution and deployment of large-scale cluster models ultimately.

[0063] S4. Inject a lightweight adaptation module at the terminal to quickly eliminate the feature differences between devices and complete the accurate adaptation of the model to the terminal environment.

[0064] The specific operation process for the lightweight adaptation module to eliminate feature differences includes: S41. Generate an adapter weight matrix based on the device feature fingerprint. For example, configure a stronger feature filtering coefficient for high-noise devices. There is a formula:

[0065] Where, is the adapter weight matrix after weight initialization, with a dimension of , represents a two-layer fully connected network with an input of the device feature vector ; S42. After the terminal loads the compressed model, freeze the parameters of the backbone network and perform 3 rounds of iterative fine-tuning on the adapter; During the iterative fine-tuning process, adjust based on the loss function of the adapter. There is:

[0066] Where, is the loss function of the adapter, is the balance coefficient, and linearly decays from 1.0 to 0.7, is the cross-entropy loss function, is the adapter weight matrix after the previous round of adjustment, initially is 0, is the square of the L2 norm; During the adjustment process, use local historical data (storage capacity < 10MB) for forward inference and loss calculation, and prevent overfitting by restricting the gradient update amplitude (learning rate ≤ 0.001); S43. Dynamically correct the feature activation threshold according to real-time environmental parameters (temperature, humidity). For example, when the temperature rises by 1°C, the voltage sag detection threshold is lowered by 0.5%. At this time, the threshold adjustment rule is as follows:

[0067] where, is the basic feature activation threshold, is the updated and adjusted feature activation threshold, is the temperature sensitivity coefficient (with a typical value ), is the nominal temperature (usually 25°C), is the actual temperature.

[0068] In a specific embodiment, a lightweight dynamic adapter (with the number of parameters <500) is injected at the input / output end of the model. The adaptation parameters are initialized through the device feature fingerprint, and the feature distribution alignment error is suppressed to less than 5% within 3 rounds of fine-tuning. The group false alarm rate is reduced from 24% to 8%. At the same time, a constrained fine-tuning algorithm is also developed to freeze the backbone network of the model and only optimize the adapter parameters, reducing the fine-tuning energy consumption of the Cortex-M7 device from 1.2J to 0.3J and compressing the peak memory occupancy to 82% of the safety threshold. Therefore, at this time, by setting the lightweight adapter injection and dynamic threshold adjustment strategy in the online fine-tuning calibration, the feature differences between devices can be quickly eliminated, the feature alignment error can be suppressed within 5%, and the group false alarm rate can be greatly reduced, facilitating the adaptation to the real-time calibration requirements in complex environments.

[0069] S5. The terminal regularly reports the operation metrics including the model inference latency, memory occupancy rate, and false alarm rate. The cloud analyzes the cluster-level data to identify devices with abnormal performance (such as the inference latency exceeding 2 times the standard deviation of the mean), and dynamically adjusts the next-round compression parameters according to the feedback data. For example, the quantization accuracy of the key layers of the device group with a high false alarm rate is increased to 6bit.

[0070] The specific operation process of dynamically adjusting the next-round compression parameters according to the feedback data includes (taking the group false alarm rate as an example): S51. Construct an iterative formula for the compression strategy, as follows:

[0071]

[0072] where, is the weight of the group false alarm rate (operation metric), is the group false alarm rate, is the average false alarm rate of the cluster, and when for the group The pruning rate of the key layer is reduced by 10%; S52. Optimize the transmission path, and use the transmission path evaluation as the optimization criterion, there is;

[0073] Among them, is the transmission path evaluation after optimization, is the bandwidth path of the th hop of the path, is the maximum bandwidth path allowed by the path, is the end-to-end delay of the th hop of the path, is the maximum allowed end-to-end delay.

[0074] At this time, by setting the performance index collection and compression strategy iteration mechanism in the closed-loop feedback optimization, the model compression parameters can be dynamically optimized to improve the quantization accuracy of the high false alarm rate device group, and achieve the dynamic balance between the model performance and the terminal resource constraints.

[0075] Therefore, in summary, this method can solve the four major problems of device heterogeneity adaptation, group transmission optimization, feature drift suppression, and resource constraint breakthrough systemically, achieve the breakthrough of the model cross-terminal migration bottleneck of heterogeneous terminal clusters, and thus provide reliable technical support for the large-scale deployment of intelligent electricity meter clusters.

[0076] In a further specific embodiment, there is a comparison of the implementation of this application with other traditional compression transfer methods. By taking the process of updating and reasoning the heterogeneous cluster model of intelligent electricity meters as the technical comparison scenario, there is: In the test environment setting, the device scale is: 1000 intelligent electricity meters (including 3 types of hardware configurations), which are respectively: High-configuration group (300 units): ARM Cortex-A53 / 1GHz, 1GB RAM, 16bit ADC Standard group (500 units): RISC-V / 800MHz, 512MB RAM, 14bit ADC Low-configuration group (200 units): Cortex-M7 / 480MHz, 256MB RAM, 12bit ADC The task goal at this time is: Deploy the ResNet-18 harmonic detection model (original size 178MB), complete the model update of the entire cluster and real-time anomaly detection.

[0077] For the comparison scheme and result analysis, there is: (1) Comparison of model compression efficiency

[0078] The traditional solution uses a unified pruning rate (50%) and 8-bit quantization for all devices, resulting in wasted computing power for high-end devices and insufficient memory for low-end devices. The present invention realizes device-level optimization by dynamically allocating the pruning rate (40% for the high-end group and 70% for the low-end group) and the quantization bit width (8 bits for the high-end group and 4 bits for the low-end group).

[0079] (2)Cluster update efficiency comparison

[0080] The traditional solution updates the model through MQTT unicast in full volume, generating 72% redundant traffic. The present invention constructs a topology routing, where the proxy nodes preferentially receive the update packets (only transmitting the differential parameters ), and distributes them through multi-hop broadcasting, increasing the transmission efficiency of the proxy nodes in the high-end group by 3 times.

[0081] (3)Detection accuracy consistency comparison

[0082] The traditional solution does not compensate for the ADC accuracy differences (12 - 16 bits) between devices, resulting in an offset in the input feature distribution. The present invention injects a dynamic adapter (with the number of parameters < 500) and adjusts the activation threshold , adaptively increasing the feature filtering intensity of high-noise devices by 30%.

[0083] (4)Edge-side resource consumption comparison

[0084] The traditional solution requires full-parameter fine-tuning (10 rounds of iteration), while the present invention constrains the fine-tuning range (only optimizing the adapter parameters), combined with a hybrid quantization strategy, reducing the inference latency of the low-end group from 189 ms to 82 ms.

[0085] In further data results, through comparison, it can be obtained that this method, compared with the traditional solution: (1)Dynamic compression gain: High-end group pruning rate optimization: (The traditional solution enforces 50%) Low-end group quantization bit width: (The traditional solution is unified at 8 bits) (2)Transmission efficiency improvement: Proxy node differential aggregation: (The total amount of differential parameters for a single device is 1080) Priority routing: (High-end group gives priority to transmission) (3)Accuracy consistency guarantee: Dynamic threshold adjustment: (Threshold is lowered by 18% at 40°C) Feature alignment error: high - configuration vs. low - configuration (After adapter compensation) In summary, this method is significantly superior to traditional solutions in core metrics such as compression ratio (+13%), transmission efficiency (-61.7% in time consumption), accuracy consistency (-66.7% in false - alarm rate), and resource economy (-75% in energy consumption), providing a feasible cross - terminal migration solution for heterogeneous terminal clusters (smart meters) in the smart grid.

[0086] On the other hand, the present invention also discloses a computer - readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above - mentioned method.

[0087] On yet another hand, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above - mentioned method.

[0088] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which, when running on a computer, causes the computer to execute any of the large - model cross - terminal transfer compression methods applied to smart electricity meters in the above - mentioned embodiments.

[0089] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. Explanations, examples, and beneficial effects of related content can refer to the corresponding parts in the above - mentioned method.

[0090] The embodiments of the present application also provide an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. The memory is used to store a computer program. The processor is used to implement the large - model cross - terminal transfer compression method applied to smart electricity meters when executing the program stored in the memory.

[0091] The communication bus mentioned in the above - mentioned electronic device can be a peripheral component interconnect standard bus or an extended industry standard architecture bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0092] The communication interface is used for communication between the above - mentioned electronic device and other devices.

[0093] The memory can include a random access memory and can also include a non - volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0094] The above-mentioned processor may be a general-purpose processor, including a central processing unit, a network processor, etc.; it may also be a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0095] It should also be noted that the electronic device further includes a terminal device, which may also be referred to as a terminal, a user equipment, a mobile station, a mobile terminal, etc. The terminal device may be a mobile phone, a smart TV, a wearable device, a tablet computer, a computer with wireless transceiver function, a virtual reality terminal device, an augmented reality terminal device, a wireless terminal in industrial control, a wireless terminal in unmanned driving, a wireless terminal in remote surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, and so on. The specific technologies and specific device forms adopted by the terminal device in the embodiments of the present application are not limited.

[0096] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server, a data center, etc. that includes one or more available media integrated. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium, or a semiconductor medium (such as a solid-state drive), etc.

[0097] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

[0098] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, then such directional indications are only used to explain the relative positional relationship, movement conditions, etc. between components in a certain specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0099] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, then such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes scenario A, or scenario B, or the scenario where both A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality of" means two or more. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

Claims

1. A large model cross - end transfer compression method applied to smart electricity meters, characterized in that, Including: S1. Quantitatively model the hardware characteristics of the terminal device, and at the same time, generate a feature fingerprint for guiding model compression based on the quantization data; S2. Dynamically adjust the compression strategy of each layer of the model structure after quantization modeling based on the fingerprint, including: using the feature fingerprint to parse and identify the computationally intensive layer, memory-sensitive layer, and precision-sensitive layer of the model structure; And sequentially perform operations: dynamically adjust the pruning rate of the computationally intensive layer according to the device computing power index, perform differential and quantization operations on the memory-sensitive layer, and perform hybrid quantization on the precision-sensitive layer to retain the FP16 precision of the key channels, ensuring that the harmonic feature extraction error is less than 0.5%; S3. Achieve efficient distribution of cluster-level model updates by constructing a device topology-aware transmission network; S4. Inject a lightweight adaptation module at the terminal to quickly eliminate the feature differences between devices and complete the precise adaptation of the model to the terminal environment.

2. The large model cross-terminal transfer compression method applied to the smart electricity meter according to claim 1, wherein During the quantization modeling process in S1, the cloud analyzes device parameters including the computing power index, memory adaptation degree, and precision correction factor to construct a multi-dimensional feature vector, complete the modeling quantization operation, and use the hierarchical clustering algorithm to divide terminals with similar features into isomorphic groups, where: The calculation formula for the computing power index is: Among them, is the main frequency of the processor of the terminal device, is the number of cores of the terminal device, is the number of instructions per cycle of the terminal device; The evaluation calculation formula for the memory adaptation degree is: Among them, is the available memory of the terminal, is the number of parameters of the benchmark model, is the total memory value of the terminal, and indicates sufficient memory, indicates insufficient memory; The calculation formula for the precision correction factor is: Among them, is the ADC root mean square noise, is the ADC resolution.

3. The large model cross-terminal transfer compression method applied to the smart electricity meter according to claim 1, characterized in that, The specific operation process of the operations performed in S2 includes: Dynamically adjust the pruning rate of the computationally intensive layer according to the device computing power index, where the pruning rate calculation formula is: Among them, represents the value of the highest computing power device in the cluster, which is the value of this computing power device, and there are constraint conditions: ; Differentiation and quantization operations on the memory-sensitive layer. At this time, the quantization bit-width allocation of the memory-sensitive layer is as follows: Among them, is the maximum memory adaptation value in the cluster, is the memory adaptation value of the computing power device, and there is a lower limit constraint: , denotes the floor symbol, which is used to represent the largest integer not greater than ; Hybridly quantize the precision-sensitive layer, with: Among them, is the quantization step, which is used to retain the FP16 precision of the key channels and forces the enabling of mixed quantization at 4. The large model cross-terminal transfer compression method applied to an intelligent electricity meter according to claim 1, wherein The specific operation process of performing model update distribution through the transmission network in S3 includes: S31. Elect a terminal with edge computing capabilities as the proxy node in each device group; S32. The cloud compares the old and new model versions, extracts the parameter change amounts of each layer, and only retains the core parameters with a change amplitude > 2%; S33. According to the geographical distribution and communication link quality of the device group, construct a multi-hop transmission path, where the proxy node preferentially receives the complete update package and then broadcasts it to the group members through the local area network, and when transmitting across groups, select the optimal path with a relay hop count ≤ 3 to avoid network congestion; S34. After the terminal receives the update package, use the hash check and rollback mechanism until the error rate < 0.01% to ensure transmission integrity.

5. The large model cross-terminal transfer compression method applied to an intelligent electricity meter according to claim 1, characterized in that The specific transmission process of the transmission network for performing efficient data transmission in S3 includes: L1. Generate a difference parameter matrix, with nodes and The matrix generation formula between them is as follows: Among them, and respectively represent the old and new model parameter matrices between nodes and ; Only the parameters are retained afterwards; L2. Calculate the transmission priority, with the calculation formula: Among them, is represented as the transmission priority, is represented as the device group the number of hops to the cloud, is represented as the device group size, is represented as the group the Frobenius norm of the parameter difference matrix; L3. The proxy node performs differential aggregation, with: Among them, is expressed as the aggregated difference parameter, is expressed as the top ones with the largest amplitudes among the selected single-device difference parameters, is expressed as the device difference parameter during the transmission process, is expressed as the device group set, is expressed as the floor symbol, used to represent the largest integer not greater than .

6. The large model cross-terminal transfer compression method applied to the smart electricity meter according to claim 1, characterized in that, The specific operation process of the lightweight adaptation module eliminating feature differences in S4 includes: S41. Generate an adapter weight matrix based on the device feature fingerprint, with the formula: Among them, is the adapter weight matrix after weight initialization, represents a two-layer fully connected network with the input being the device feature vector ; S42. After the terminal loads the compressed model, freeze the parameters of the backbone network and iteratively adjust the adapter; During the adjustment process, adjust based on the loss function of the adapter, with: Among them, is the loss function of the adapter, is the balance coefficient, and linearly decays from 1.0 to 0.7, is the cross-entropy loss function, is the adapter weight matrix after the previous round of adjustment, with the initial being 0, being the square of the L2 norm; During the adjustment process, use local historical data for forward inference and loss calculation, and prevent overfitting by restricting the gradient update amplitude; S43. Dynamically correct the feature activation threshold according to the real-time environment parameters, where the threshold adjustment rule is: Among them, is the basic feature activation threshold, is the feature activation threshold after update and adjustment, is the temperature sensitivity coefficient, is the nominal temperature, is the implementation temperature.

7. The large model cross-terminal transfer compression method applied to the smart electricity meter according to claim 1, characterized in that, It further includes S5: Regularly report operating metrics including model inference latency, memory occupancy rate, and false alarm rate through the terminal, analyze cluster-level data through the cloud, identify devices with abnormal performance, and dynamically adjust the next-round compression parameters according to the feedback data.

8. The large model cross-terminal transfer compression method applied to the intelligent electricity meter according to claim 7, characterized in that, The specific operation process of dynamically adjusting the next-round compression parameters according to the feedback data in the above S5 includes: S51. Construct an iterative formula for the compression strategy, as follows: Among them, is the weight of the operation index; S52. Optimize the transmission path, using the transmission path evaluation as the optimization criterion, as follows: Among them, is the optimized transmission path evaluation, is the bandwidth path of the th hop of the path, is the maximum bandwidth path allowed by the path, is the end-to-end delay of the th hop of the path, is the maximum allowed end-to-end delay.

9. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Topology identification model compression method, system and equipment based on fusion terminal

    CN119005263A

  • Industrial equipment real-time monitoring system based on edge computing

    CN119644972A

  • Small-amount data fine tuning and self-repairing generative model compression method for specific task of power system

    CN119830968A

  • Method and system for fine tuning and lightweight design of electric power visual large model

    CN120046683A

  • Method and apparatus for compressing topology recognition model, electronic device, and medium

    US20250086460A1

Cited By

  • Data acquisition interaction method of electric energy meter data terminal

    CN120915442A

  • A data acquisition interaction method of an electric energy meter data terminal

    CN120915442B