Artificial intelligence model self-adaptive compression method and system for power intelligent terminal
Patent Information
- Application Number
- CN202611281855.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-25
AI Technical Summary
[0011]本发明所要解决的技术问题在于针对上述现有技术中的不足,提供一种面向电力智能终端的人工智能模型自适应压缩方法及系统,通过综合考虑模型架构、硬件环境进行深层次协同优化,用于解决现有人工智能模型压缩方法未能统一考虑电力智能终端资源约束和电力业务要求、固定压缩方案难以适配不同电力智能终端,以及在多种压缩策略组合构成的决策空间中难以兼顾搜索范围和解集分布的技术问题,有效融合用户需求与平台资源约束,实现压缩策略的自动搜索、性能优化提升
一种面向电力智能终端的人工智能模型自适应压缩方法,通过将电力AI模型的网络架构参数、电力智能终端的硬件资源条件以及电力业务需求共同引入模型压缩过程,构建以压缩策略组合为决策变量,并以推理精度、计算量、存储、时延和功耗为评价维度的多目标优化模型,使压缩策略的搜索同时受到计算资源配额、内存配额、功耗配额、推理精度阈值和时延阈值的约束,从而降低压缩后模型无法满足终端部署条件的风险。进一步地,通过方向向量将初始种群划分为多个子种群,使不同子种群分别对应不同的目标搜索方向;利用深度Q网络根据电力AI模型的网络架构参数及当前压缩策略组合关联的方向向量,对各子种群进行评价并选择子种群组成繁衍池;再通过交叉、变异和环境选择迭代更新种群,扩展候选压缩策略组合并筛除不满足约束条件的方案。同时,根据模型精度增益项和资源约束增益项计算奖励,并据此更新深度Q网络和经验池,使后续搜索能够结合已有搜索结果调整子种群选择方向。由此,实现模型压缩策略的自适应搜索,并得到满足终端资源约束和电力业务需求的轻量化模型。
Smart Images

Figure CN122819346A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power edge intelligence and artificial intelligence model compression technology, specifically relating to an adaptive compression method and system for artificial intelligence models for power intelligent terminals. Background Technology
[0002] With the development of new power systems and the Internet of Things (IoT) in the power sector, artificial intelligence (AI) models are gradually being deployed in power smart terminals such as distribution transformer convergence terminals, energy controllers, smart energy units, edge IoT agents, charging piles, and smart meters. These models enable power operations such as load forecasting, intelligent inspection, online monitoring, fault diagnosis, and demand response to be performed close to the data source. By completing data processing and inference at the power smart terminal side, the communication pressure caused by transmitting business data to the cloud can be reduced, and the real-time processing requirements of some power services can be met.
[0003] Power smart terminals typically possess data acquisition, network communication, data storage, and computational analysis capabilities. However, different power smart terminals vary in hardware conditions such as processor type, computing power, storage space, and allowable power consumption. Directly deploying a trained artificial intelligence model to a power smart terminal may result in computational load, storage consumption, inference latency, and power consumption exceeding the terminal's available resources. Therefore, before deploying an artificial intelligence model to a power smart terminal, it is usually necessary to employ model compression techniques such as model pruning, parameter quantization, low-rank decomposition, or knowledge distillation to lightweight the artificial intelligence model.
[0004] Existing model pruning methods reduce model size by removing some weights, convolutional kernels, channels, or network layers; parameter quantization methods reduce storage and computational overhead by decreasing the representation precision of model weights or activation values; low-rank decomposition methods reduce the number of model parameters by decomposing high-dimensional parameters; and knowledge distillation methods obtain lightweight models by having student models learn the output or intermediate features of teacher models. These methods typically require pre-selection of compression techniques, compression ratios, and corresponding hyperparameters. Different compression techniques and their combinations have varying impacts on model accuracy, computational cost, storage, latency, and power consumption.
[0005] An existing model compression method combines pruning and quantization. First, the model is structurally pruned based on the information complexity and contribution of the convolutional kernels. Then, the weights of the pruned model are quantized. This method can reduce the number of model parameters and computational overhead. However, its pruning rules, quantization methods, and execution order all need to be pre-set, and the generated compression scheme is mainly targeted at specific models and hardware environments. When the terminal processor type, computing power, storage space, or power consumption conditions change, the compression parameters or compression process need to be readjusted, making it difficult to directly adapt to power smart terminals with different hardware configurations. Furthermore, this method does not uniformly model inference accuracy, computational load, storage, latency, and power consumption with terminal resource quotas, which may result in situations where the compressed model accuracy meets requirements, but the computational load, storage usage, inference latency, or power consumption still exceeds the terminal resource limits.
[0006] Another existing automated model compression method based on Q-Learning sets information such as model inference time, model size, power consumption, and accuracy as states, and predefined pruning, quantization, and sparsification operations as actions, selecting a model compression scheme through reinforcement learning. This method reduces the workload of manually selecting compression methods, but its state space and action space consist of a finite number of pre-defined compression operations, limiting the combinations of compression strategies that can be searched. When the number of compression strategy combinations increases or multiple resource constraints need to be considered simultaneously, a single reinforcement learning search process struggles to balance search scope and efficiency. Furthermore, this method does not consider the computing resource quota, memory quota, and power consumption quota of the power smart terminal as mandatory constraints, meaning the obtained compression scheme may not be deployable on power smart terminals with strictly limited resources.
[0007] Therefore, it is evident that existing artificial intelligence model compression technologies suffer from at least the following problems: First, existing model compression methods mainly select compression schemes based on model structure or model accuracy, without uniformly modeling the computing resources, storage resources, and power consumption conditions of power smart terminals with the inference accuracy and time delay requirements of power services. This makes it difficult to ensure that the compressed model simultaneously meets the model performance requirements and terminal resource constraints.
[0008] Secondly, existing compression processes typically rely on pre-set compression technologies, compression sequences, and compression parameters. Compression schemes obtained for one hardware platform are difficult to directly apply to other smart power terminals with different hardware configurations, resulting in insufficient cross-terminal platform adaptability.
[0009] Third, the search space of existing automated model compression methods is usually composed of limited preset actions, making it difficult to effectively search in a large decision space that includes multiple compression techniques and their parameter combinations. At the same time, it is difficult to take into account the diversity and convergence between different search directions.
[0010] Fourth, existing methods lack a feedback mechanism that combines the compression strategy search process with model accuracy gains and resource constraint improvements, making it difficult to adaptively adjust the compression strategy search direction based on the network architecture parameters of different power artificial intelligence (AI) models and the search status of the current compression strategy combination. Summary of the Invention
[0011] The technical problem to be solved by this invention is to address the shortcomings of the prior art by providing an adaptive compression method and system for artificial intelligence models for power smart terminals. By comprehensively considering the model architecture and hardware environment for deep collaborative optimization, this invention solves the technical problems of existing artificial intelligence model compression methods failing to uniformly consider the resource constraints and power business requirements of power smart terminals, the difficulty of adapting fixed compression schemes to different power smart terminals, and the difficulty of balancing the search range and solution set distribution in the decision space composed of multiple compression strategy combinations. This invention effectively integrates user needs and platform resource constraints, and achieves automatic search of compression strategies and performance optimization.
[0012] The present invention adopts the following technical solution: Firstly, an adaptive compression method for artificial intelligence models in power smart terminals is provided, which is applied to power smart terminals and includes the following steps: S1. Perform network architecture analysis on the trained power AI model to obtain the model structure features; S2. Perform resource quota mapping on the underlying hardware parameters and system operation status of the power intelligent terminal to obtain a terminal resource quota set; S3. Constraint modeling is performed on the power business demand and the set of terminal resource quotas to obtain a multi-objective optimization model with the combination of compression strategies as decision variables; S4. Generate direction vectors in the target space of the multi-objective optimization model, and construct a population of compression strategy combinations to obtain a set of direction vectors and an initial population. S5. Perform directional vector association partitioning on the compression strategy combinations in the initial population to obtain multiple subpopulations; S6. The current population state, which is composed of the structural features of the model and the direction vectors associated with each subpopulation, is evaluated by a deep Q-network to obtain the Q value of each subpopulation, and a breeding pool is obtained by screening based on the Q value. S7. Perform crossover and mutation on the compression strategy combinations in the breeding pool to generate offspring solutions, and perform environmental selection on the parent population and the offspring solutions to obtain an updated population. S8. Calculate the reward for the performance index of the updated population to obtain the reward value, and update the deep Q network and experience pool according to the reward value. S9. Repeat S6 to S8 until the termination condition is met to obtain the optimal compression strategy combination, and compress the power AI model based on the optimal compression strategy combination to obtain a lightweight model.
[0013] Furthermore, the power AI model includes a new energy output prediction model, a load prediction model, an electricity consumption behavior analysis model, an intelligent inspection model, an online monitoring model, a fault judgment model, a distribution network topology identification model, or a regulation strategy generation model.
[0014] Furthermore, the model structure features include at least one of the following: weight parameters, activation value parameters, layer type, and number of layers in the power AI model.
[0015] Furthermore, in S2, the underlying hardware parameters are obtained by calling the hardware abstraction layer interface of the power smart terminal, and the resource quota, memory quota, and power consumption quota are set according to the underlying hardware parameters to obtain the terminal resource quota set.
[0016] Furthermore, in S3, the power service requirements include the inference accuracy threshold and latency threshold corresponding to the power service, and the terminal resource quota set includes computing resource quota, memory quota and power consumption quota; For any of the compression strategy combinations, the objective space of the multi-objective optimization model includes the inference accuracy, computational load, storage, latency, and power consumption of the power AI model after processing by the compression strategy combination. The constraints of the multi-objective optimization model include: the inference accuracy is not lower than the inference accuracy threshold, the computational load does not exceed the computational resource quota, the storage does not exceed the memory quota, the latency does not exceed the latency threshold, and the power consumption does not exceed the power consumption quota. The inference accuracy is: the ratio of the number of samples whose inferred values are equal to the true values to the total number of samples in the validation set when the power AI model after the compression strategy is combined and processed to infer the validation set samples. The computational load is determined based on the total number of floating-point arithmetic operations performed by the power AI model in one forward propagation; The storage is determined based on the model parameters, activation values, and total memory required by the runtime environment of the power AI model; The delay is determined based on the sum of the computation time, data and instruction reading time, and scheduling time required for the power AI model to go from receiving input to outputting the final result; The power consumption is determined based on the sum of the power consumption calculated by the model and the power consumption for memory read / write operations.
[0017] Furthermore, S5 includes: Calculate the angle between each compression strategy combination in the initial population and each direction vector in the direction vector set; divide each compression strategy combination into a subpopulation corresponding to the direction vector with the smallest angle.
[0018] Further, in S6: the state of the deep Q-network includes the model structure features and the direction vector associated with the current compression strategy combination to be evaluated; the action of the deep Q-network includes selecting one subpopulation from the multiple subpopulations. In S8, the reward value is determined based on the model accuracy gain term and the resource constraint gain term. The model accuracy gain term is used to represent the change in inference accuracy of the child solution relative to the parent solution, and the resource constraint gain term is used to represent the change in computation, storage, latency and power consumption of the child solution relative to the parent solution. The model accuracy gain term and the resource constraint gain term are normalized and then combined according to their corresponding weight coefficients, which are determined by the fitness function.
[0019] Furthermore, in S7: Two parent individuals are selected from the current population using either sorting selection or Softmax probabilistic selection. Then, the compression strategy of the two parent individuals is combined using arithmetic crossover or heuristic crossover to generate offspring individuals. Gaussian noise is added to the original values using Gaussian mutation to alter at least one gene in the offspring individuals.
[0020] Furthermore, the environment selection strategy is a non-dominated sorting-based environment selection strategy. In S7, environment selection is performed on the parent population and the child solutions to obtain an updated population, including: The parent population and the offspring population are merged to form a temporary population; The temporary population is subjected to fast non-dominated sorting, and the crowding distance of each individual in each non-dominated layer is calculated. According to the non-dominated sorting priority from high to low and the crowding distance from large to small, a preset number of individuals are selected from the temporary population to form the updated population.
[0021] Furthermore, the compression strategy combination includes at least one compression technique selected from weight pruning, quantization, sparse decomposition, and special architecture layer replacement.
[0022] Furthermore, the termination condition includes reaching the maximum number of iterations, or the improvement in inference accuracy corresponding to the optimal compression strategy combination generated in consecutive P generations being less than a preset improvement threshold; where P is a preset positive integer, and the improvement in inference accuracy is the positive increment of the inference accuracy corresponding to the current generation's optimal compression strategy combination relative to the inference accuracy corresponding to the previous generation's optimal compression strategy combination.
[0023] Secondly, embodiments of the present invention provide an adaptive compression system for artificial intelligence models of power smart terminals, used to execute the adaptive compression method for artificial intelligence models of power smart terminals described in the first aspect, the system comprising: Architecture analysis module, Used to perform network architecture analysis on trained power AI models to obtain model structural features; The resource quota mapping module is used to map the underlying hardware parameters and operating status of the power smart terminal to obtain a set of terminal resource quotas. The constraint modeling module is used to perform constraint modeling on the power business demand and the terminal resource quota set to obtain a multi-objective optimization model with the combination of compression strategies as decision variables. The model compression module is used to search for compression strategy combinations based on the multi-objective optimization model, the model structural features, and the terminal resource quota set to obtain the optimal compression strategy combination, and to compress the power AI model based on the optimal compression strategy combination to obtain a lightweight model.
[0024] Furthermore, the model compression module is used for: Direction vectors are generated in the target space of the multi-objective optimization model, and a population is constructed for the combination of compression strategies to obtain a set of direction vectors and an initial population. The compression strategy combinations in the initial population are partitioned by directional vector association to obtain multiple subpopulations; The current population state, which is composed of the structural features of the model and the direction vectors associated with each subpopulation, is evaluated by a deep Q-network to obtain the Q value of each subpopulation, and a breeding pool is obtained based on the Q value. Crossover and mutation are performed on the compression strategy combinations in the breeding pool to obtain offspring solutions, and environmental selection is performed on the parent population and the offspring solutions to obtain an updated population. The performance metrics of the updated population are used to calculate rewards, and the reward value is then used to update the deep Q-network and the experience pool. The deep Q-network evaluation, crossover and mutation, environment selection, and reward calculation are performed iteratively until the termination condition is met, thus obtaining the optimal compression strategy combination.
[0025] Thirdly, a smart power terminal includes a processor and a memory; The memory is used to store multiple programs, including an adaptive compression system for artificial intelligence models for smart power terminals; When the processor executes the AI model adaptive compression system for power smart terminals, it causes the processor to execute the AI model adaptive compression method for power smart terminals as described in the first aspect.
[0026] Furthermore, the processor includes at least one of a central processing unit, a graphics processing unit, a neural network processing unit, or a field-programmable gate array; The memory also stores the operating system of the power smart terminal. The operating system includes a kernel layer, a system service layer, and an application layer. When the computer program is executed by the processor, it obtains the underlying hardware parameters of the power smart terminal through the interface of the operating system and deploys the compressed lightweight model on the system service layer for edge computing services.
[0027] Fourthly, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the adaptive compression method for an artificial intelligence model for a power smart terminal described in the first aspect.
[0028] Compared with the prior art, the present invention has at least the following beneficial effects: An adaptive compression method for artificial intelligence models in power smart terminals is proposed. This method incorporates the network architecture parameters of the power AI model, the hardware resources of the power smart terminal, and the power business requirements into the model compression process. A multi-objective optimization model is constructed, using compression strategy combinations as decision variables and inference accuracy, computational load, storage, latency, and power consumption as evaluation dimensions. This ensures that the search for compression strategies is simultaneously constrained by computational resource quotas, memory quotas, power consumption quotas, inference accuracy thresholds, and latency thresholds, thereby reducing the risk that the compressed model will not meet the terminal deployment requirements. Furthermore, the initial population is divided into multiple subpopulations using direction vectors, with each subpopulation corresponding to a different target search direction. A deep Q-network is used to evaluate each subpopulation based on the network architecture parameters of the power AI model and the direction vector associated with the current compression strategy combination, and selects subpopulations to form a breeding pool. The population is then iteratively updated through crossover, mutation, and environmental selection to expand candidate compression strategy combinations and eliminate schemes that do not meet the constraints. Simultaneously, rewards are calculated based on model accuracy gain and resource constraint gain, and the deep Q-network and experience pool are updated accordingly, enabling subsequent searches to adjust the subpopulation selection direction based on existing search results. Thus, an adaptive search for model compression strategies is achieved, resulting in a lightweight model that meets terminal resource constraints and power business requirements.
[0029] Furthermore, the type of power AI model is limited, enabling this invention to be applied to various power services such as renewable energy output forecasting, load forecasting, electricity consumption behavior analysis, intelligent inspection, online monitoring, fault diagnosis, distribution network topology identification, and regulation strategy generation. By limiting the compression object to power AI models with clear business applications, it is easier to generate corresponding compression strategy combinations based on the structural characteristics and deployment requirements of different business models.
[0030] Furthermore, by limiting the network architecture parameters and analyzing the weight parameters, activation value parameters, layer type, and number of layers of the power AI model, the deep Q-network can obtain state information reflecting the current model structure characteristics. Since different model structures have varying degrees of applicability to pruning, quantization, and structural replacement, this facilitates ensuring that the subpopulation selection process corresponds to the actual architecture of the model to be compressed.
[0031] Furthermore, by calling the hardware abstraction layer interface of the power intelligent terminal operating system to obtain underlying hardware parameters, the model compression process can set resource quotas based on the actual hardware configuration of the terminal and the system's operating status. The hardware abstraction layer interface can provide a relatively unified entry point for obtaining hardware information, which helps to reduce the impact of differences in hardware access methods between different processors, memory, and terminal platforms on the resource perception process.
[0032] Furthermore, the methods for determining inference accuracy, computational load, storage, latency, and power consumption are clearly defined, providing a clear computational basis for each evaluation index in the multi-objective optimization model. By evaluating compression strategy combinations from aspects such as model output accuracy, floating-point operation load, memory usage, inference time composition, and computation and memory access power consumption, the bias in strategy comparison caused by inconsistent evaluation criteria is reduced.
[0033] Furthermore, by calculating the angle between the compression strategy combination and the direction vector, and dividing the compression strategy combination into the subgroup corresponding to the direction vector with the smallest angle, each compression strategy combination can be classified according to its direction in the target space. This division method helps to maintain different target search directions and avoids candidate solutions from being overly concentrated on a single performance index.
[0034] Furthermore, the states, actions, and rewards of the deep Q-network are constrained to provide a clear decision-making basis for the subpopulation selection process. By using the model network architecture parameters and associated direction vectors as states, subpopulation selection as actions, and calculating rewards based on model accuracy gain terms and resource constraint gain terms, the search process simultaneously considers changes in model performance and resource constraints.
[0035] Furthermore, by selecting parent individuals and performing arithmetic crossover, heuristic crossover, and Gaussian mutation, existing compression strategies are recombined and locally modified. The crossover operation is used to inherit effective components from different parent compression strategies, while the mutation operation is used to introduce new candidate solutions, expand the search range of compression strategy combinations, and reduce the possibility of the search process getting stuck in local regions.
[0036] Furthermore, an environmental selection strategy based on non-dominated ordination and crowding distance is adopted to uniformly screen both parent and offspring populations. Non-dominated ordination is used to distinguish the priority of different candidate schemes on multiple evaluation indicators, while crowding distance is used to evaluate the distribution of candidate schemes in the target space, taking into account both constraint compliance and scheme distribution when updating the population.
[0037] Furthermore, the compression strategy combination is limited to at least one compression technique among weight pruning, quantization, sparse decomposition, and special architecture layer replacement, enabling the search object to cover different types of model compression methods. By selecting or combining multiple compression techniques individually, different compression paths are formed according to the model structure and terminal resource conditions, reducing the limitation of always using a single compression method.
[0038] Furthermore, the search process is controlled by setting a termination condition based on either the maximum number of iterations or the inference accuracy improvement corresponding to a combination of consecutive generations of compression strategies. The maximum number of iterations limits the computational cost of the search process, while the termination condition corresponding to the inference accuracy improvement ends the search when continuous iterative improvement is finite, avoiding ineffective repeated iterations.
[0039] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0040] In summary, this invention integrates the power AI model structure, terminal hardware resources, and business requirements to automatically search for compression strategy combinations that meet the constraints of accuracy, computational load, storage, latency, and power consumption. This reduces manual parameter tuning, improves the cross-terminal adaptability of model compression, and yields a lightweight model suitable for deployment on power smart terminals.
[0041] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the steps of the method of the present invention; Figure 2 This is a schematic diagram of the system architecture of the present invention; Figure 3 This is a schematic diagram of the structure of the power intelligent terminal of the present invention; Figure 4This is a schematic diagram of the structure of the power intelligent terminal operating system of the present invention; Figure 5 A schematic diagram of a computer device provided in an embodiment of the present invention; Figure 6 This is a block diagram of a chip provided according to an embodiment of the present invention.
[0043] Among them, 60. Computer equipment; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access memory unit; 6202. Cache memory unit; 6203. Read-only memory unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0046] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0047] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.
[0048] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0049] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0050] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.
[0051] This invention provides an adaptive compression method for artificial intelligence models of power smart terminals. Based on multi-swarm co-evolutionary algorithms and deep reinforcement learning, it achieves AI model compression and optimization for power smart terminals. By actively sensing the hardware configuration of power smart terminals, it uses a combination of reinforcement learning and multi-swarm co-evolutionary algorithms to search for optimal hyperparameters under hardware constraints, enabling the model to adapt to different power smart terminals. This addresses the problems of existing model compression methods failing to uniformly model the hardware resource constraints of power smart terminals with power business requirements, relying on preset rules for compression strategies that are difficult to adapt to different power smart terminals, and struggling to balance search range and search efficiency in a decision space containing multiple compression strategy combinations.
[0052] Please see Figure 1This invention discloses an adaptive compression method for artificial intelligence models of power intelligent terminals. It acquires the network architecture parameters of a trained power AI model and the underlying hardware parameters of the power intelligent terminal, sets computational resource quotas, memory quotas, and power consumption quotas, and combines these with inference accuracy and latency thresholds corresponding to power services to construct a multi-objective optimization model with compression strategy combinations as decision variables. Direction vectors are generated in the objective space of the multi-objective optimization model, and the initial population composed of multiple compression strategy combinations is divided into multiple sub-populations. A deep Q-network is used to evaluate each sub-population and select sub-populations to form a breeding pool. The population is then updated through crossover, mutation, and environmental selection. Rewards are calculated based on model accuracy gain and resource constraint gain, and the deep Q-network and experience pool are updated based on these rewards until the optimal compression strategy combination that satisfies the constraints of the multi-objective optimization model is output. This process compresses the power AI model to obtain a lightweight model. The specific steps are as follows: S1. Perform architecture analysis on the trained power AI model to obtain the model's structural features; The power AI model can be deployed from the cloud or obtained from a local AI model library.
[0053] The power AI model includes: a new energy output prediction model, a load prediction model, an electricity consumption behavior analysis model, an intelligent inspection model, an online monitoring model, a fault analysis model, a distribution network topology identification model, and a regulation strategy generation model.
[0054] In this embodiment, an online image monitoring model for an anomaly identification scenario in a power distribution station is taken as an example. The trained online image monitoring model is obtained, and its network architecture parameters are analyzed. These parameters include at least one of model weight parameters, activation value parameters, layer type, and number of layers.
[0055] The power AI model to be compressed and its network structure information are used as the basic input for subsequent compression strategy search. Different network layers, channels, and convolutional kernels have different effects on the number of model parameters, computational cost, and inference accuracy. By analyzing the network architecture parameters, the subsequent deep Q-network can consider the actual structural characteristics of the current power AI model when selecting sub-populations, avoiding processing power AI models with different structures only according to fixed compression rules.
[0056] The model structure features include the weight parameters, activation value parameters, layer type and number of layers, and model type information of the power AI model, which are used as inputs for the search of the deep Q network state and compression strategy.
[0057] S2. Perform resource quota mapping on the underlying hardware parameters and system operation status of the power intelligent terminal to obtain a terminal resource quota set; Specifically, by calling the hardware abstraction layer interface of the power intelligent terminal operating system, underlying hardware parameters such as processor type, processor computing power, available memory, and allowed power consumption are obtained. Based on these underlying hardware parameters and the system's operating status, the corresponding computing resource quota, memory quota, and power consumption quota for the power AI model when running on the power intelligent terminal are set.
[0058] The computing resource quota is the upper limit of computing power minus the computing power required by the operating system, basic APP, business APP, and security reserved computing power, used to limit the upper limit of the computing amount of the compressed power AI model in one forward propagation process; the memory quota is the upper limit of memory minus the memory occupied by the operating system, basic APP, business APP, and security reserved memory, used to limit the upper limit of the storage space occupied by model parameters, activation values, and runtime environment; the power consumption quota is the upper limit of power consumption at the current temperature minus the power consumption of the operating system, basic APP, business APP, and security reserved power consumption, used to limit the upper limit of power consumption of the compressed power AI model in the inference process.
[0059] The basic APP is a system-level application necessary for the operation of the power smart terminal, used to ensure the terminal's basic functions, communication interaction, safe operation, and data management; the business APP is an application deployed according to the specific power business needs such as power distribution, marketing, and dispatch.
[0060] The security-reserved computing power is a portion of computing resources reserved in advance to ensure system anomaly handling and sudden business needs; the security-reserved memory is a portion of memory resources reserved in advance to ensure system anomaly handling and sudden business needs; the security-reserved power consumption is a power budget reserved in advance for terminal power management and temperature rise control.
[0061] The hardware capabilities of different power smart terminals are transformed into constraint boundaries that can be used by a multi-objective optimization model. Due to differences in processor, memory, and power consumption among different power smart terminals, the same compression strategy combination may be suitable for one power smart terminal but not another. By setting resource quotas for the current power smart terminal, the search results of compression strategy combinations can be made to correspond to the current hardware platform.
[0062] S3. Constraint modeling is performed on the power business demand and the set of terminal resource quotas to obtain a multi-objective optimization model with the combination of compression strategies as decision variables; The inference accuracy threshold is set according to the power business's requirements for model inference accuracy, and is used to represent the minimum inference accuracy that the compressed power AI model should achieve; the latency threshold is set according to the power business's requirements for real-time performance, and is used to represent the maximum inference latency allowed by the compressed power AI model.
[0063] In this embodiment, the inference accuracy threshold is set to 95% for the substation anomaly identification service. The latency threshold is determined based on the real-time processing requirements of the substation anomaly identification service.
[0064] Resource perception and business requirement modeling for power intelligent terminals include: acquiring the underlying hardware parameters of the power intelligent terminal; setting computing resource quotas, memory quotas, and power consumption quotas based on the underlying hardware parameters and system operating conditions; acquiring inference accuracy thresholds and latency thresholds corresponding to power services; and constructing a multi-objective optimization model corresponding to a combination of model compression strategies based on the computing resource quotas, memory quotas, power consumption quotas, inference accuracy thresholds, and latency thresholds. The inference accuracy threshold is used to limit the minimum inference accuracy that the compressed power AI model should achieve, and the latency threshold is used to limit the maximum allowable inference latency of the compressed power AI model.
[0065] The objective space of the multi-objective optimization model includes inference accuracy, computational cost, storage, latency, and power consumption. The constraints of the multi-objective optimization model include: The inference accuracy is not lower than the inference accuracy threshold, the computation amount does not exceed the computing resource quota, the storage does not exceed the memory quota, the latency does not exceed the latency threshold, and the power consumption does not exceed the power consumption quota; Define inference precision ,in, Indicates the inference value. Represents the true value. The inference precision is defined as the ratio of the inferred value to the true value, where M represents the number of samples in the validation set; No less than 95%; Wherein, inference accuracy represents the correctness of the output of the compressed power AI model on the validation set; M represents the number of validation set samples; the inference value of the i-th sample represents the prediction result output by the power AI model, and the true value represents the labeled result corresponding to the validation set sample.
[0066] Define computational quantity This represents the sum of all floating-point arithmetic operations that the power AI model, after being processed by the current compression strategy, theoretically needs to perform in a single forward propagation. ,in, Indicates the number of model layers. Indicates the first The number of floating-point operations per layer For the first power AI model The number of floating-point operations a layer needs to perform during one forward propagation; Wherein, computational complexity represents the floating-point operations required for the compressed power AI model to complete one forward propagation; model layer number represents the number of computable network layers contained in the power AI model; and the number of floating-point operations of layer l represents the number of floating-point arithmetic operations such as multiplication and addition performed by layer l in one forward propagation.
[0067] Define computing resource quotas to limit the amount of computing power. The upper limit; Define storage This represents the total memory required for the power AI model after processing using the current compression strategy, including the parameters, activation values, and runtime environment. ,in, For the number of model parameters, For intermediate activation value memory, This refers to the runtime memory overhead of the AI framework. Among them, storage represents the memory required for the deployment and operation of the compressed power AI model; model parameter quantity represents the memory occupied by learnable parameters such as weights and biases; intermediate activation value memory represents the temporary memory occupied by the output features of each layer during model inference; and AI framework runtime memory overhead represents the memory required for loading and scheduling the model runtime environment.
[0068] Define storage Memory quota for power smart terminals; Define delay This represents the total time taken by the power AI model, after processing using the current compression strategy, from receiving input to outputting the final result. ,in, To calculate time, For data and instruction read time, For scheduling time; Among them, latency represents the time required for the compressed power AI model to complete one end-to-end inference; computation time represents the execution time of the model operator; data and instruction read time represents the time required to read input data, model parameters and execution instructions from memory to the processing unit; and scheduling time represents the time generated by the operating system or runtime framework for task scheduling.
[0069] Define delay Business requirements for model inference in specific business scenarios; Define the power consumption of model inference This is the sum of power consumption and memory read / write power consumption calculated for the power AI model after processing with the current compression strategy. ,in, Calculate power consumption for the model. Power consumption for memory read / write; Among them, power consumption represents the electrical power consumed by the compressed power AI model during inference; model computation power consumption represents the power consumption generated by the processor executing model operators; memory read / write power consumption represents the power consumption generated when reading model parameters, input data and intermediate activation values and writing output results.
[0070] Define power consumption Power consumption quotas for smart power terminals.
[0071] The business requirements are modeled as follows:
[0072] ,
[0073]
[0074] in, For a combination of compression strategies, The decision space constitutes all possible compression strategies. Let be the objective function. The negative inference accuracy of the power AI model after processing with the compression strategy combination x is given. This represents the computational cost of the power AI model processed by the compression strategy combination x during one forward propagation process. The storage footprint of the power AI model after processing with compression strategy combination x. The inference latency of the power AI model after processing with the compression strategy combination x is given. The inference power consumption of the power AI model after processing with the compression strategy combination x is given. This is the inference accuracy threshold. To calculate resource quotas, For memory quota, For the time delay threshold, This refers to power consumption quotas.
[0075] Wherein, the compression strategy combination represents a candidate scheme consisting of at least one compression technique and its parameters, which is selected from weight pruning, quantization, sparse decomposition or special architecture layer replacement; the decision space represents the set of all candidate compression strategy combinations; and the objective function represents an index vector consisting of inference accuracy, computational cost, storage, latency and power consumption used to evaluate the candidate scheme.
[0076] The model inference performance, power smart terminal resource conditions, and power business requirements are uniformly transformed into evaluation indicators and constraints within a single multi-objective optimization model. Therefore, candidate compression strategy combinations not only need to maintain inference accuracy but also simultaneously meet limitations in computation, storage, latency, and power consumption, reducing situations where optimizing only model accuracy or compression rate results in lightweight models failing to run on power smart terminals.
[0077] S4. Generate direction vectors in the target space of the multi-objective optimization model, and generate an initial population including combinations of multiple compression strategies; initialize the direction vectors and generate a set of uniformly distributed vectors in the target space. A dimensional direction vector is used to guide the search direction; In this embodiment, the direction vector is in the target space. M dimensional vector, M The target quantity. In this embodiment, the targets include inference accuracy, computational load, storage, latency, and power consumption. The set of direction vectors includes K A directional vector, K As a preset positive integer, K The value of is set based on the population size and the precision of the target space partitioning. Each direction vector corresponds to a subpopulation; therefore, the number of subpopulations is the same as the number of direction vectors. Direction vectors can be obtained by normalizing... M The target space is uniformly sampled to generate different directional vectors, so that different directional vectors correspond to different target search directions.
[0078] Here, the direction vector represents the search direction in the target space; the target space dimension corresponds to five evaluation dimensions: inference accuracy, computational cost, storage, latency, and power consumption; and the initial population represents the set of candidate solutions composed of multiple candidate compression strategies.
[0079] Initialize the population and generate a population containing Initial population of candidate compression technology combinations .
[0080] The compression techniques include weight pruning, quantization, sparse decomposition, and replacement of special architecture layers.
[0081] Multiple search directions are provided in a search space with multiple mutually constraining objectives. Since increasing the compression ratio may lead to a decrease in inference accuracy, and there may be mutual influences between reducing computation, storage, latency, and power consumption, by setting multiple direction vectors, different combinations of candidate compression strategies can be searched along different target directions.
[0082] S5. Perform directional vector association partitioning on the compression strategy combinations in the initial population to obtain multiple subpopulations; Divide the population and divide the initial population Associating uniformly distributed direction vectors Calculate the angle between each individual in the population and the direction vector, and assign each individual to the subpopulation corresponding to the direction vector with the smallest angle. Each direction vector corresponds to one subpopulation, and the initial population is decomposed into multiple subpopulations. Each subpopulation contains all those associated with Each subpopulation represents a specific target search direction.
[0083] Wherein, the included angle represents the directional deviation between the objective function vector and the direction vector corresponding to the compression strategy combination; the smallest included angle indicates that the compression strategy combination is closest to the search direction guided by the corresponding direction vector in the target space.
[0084] The initial population is divided according to the orientation of the compression strategy combination in the target space, so that different subpopulations correspond to different target trade-offs. This division method can reduce the over-concentration of candidate compression strategy combinations in a single target direction and enable the subsequent deep Q-network to evaluate different search directions based on the subpopulations.
[0085] S6. Input the current population state into the deep Q network, evaluate the Q value of each subpopulation, and select the subpopulations to form a breeding pool. The selection of potential subpopulations is guided by a deep Q-network, based on the current population state. Input DQN to calculate the values of each subpopulation. Value, selection Subpopulations with higher values form a breeding pool to ensure the selection of high-potential compression combinations. The deep Q-network includes: Define state For the network architecture parameters of the AI model currently being optimized and the direction vector associated with the current solution, formally, ; Here, the state represents the decision input received by the deep Q-network in the current iteration; the network architecture parameters are used to characterize the structural features of the power AI model to be compressed; and the direction vector is used to characterize the search direction of the current compression strategy combination in the target space.
[0086] Define Action To select from all candidate subpopulations Select a specific subpopulation from the population; Here, "action" represents the subpopulation selection result output by the deep Q-network; "candidate subpopulation" represents the set of compression strategy combinations obtained by partitioning the direction vector.
[0087] Define rewards The benefits of improving the parent solution by the child solution are as follows:
[0088]
[0089]
[0090] in, This is the model accuracy gain term. For resource constraint gain term, Represents the normalization function. These are the weighting coefficients, which are determined by the fitness function. As a reward, For computational load, For the sake of reasoning accuracy, This is the inference accuracy threshold. For storage, For time delay, For power consumption, To calculate resource quotas, For memory quota, For the time delay threshold, For power consumption quota; Among them, the model accuracy gain term represents the change in inference accuracy of the child solution relative to the parent solution; the resource constraint gain term represents the change in computational cost, storage, latency, and power consumption of the child solution relative to the parent solution; the normalization function is used to map indicators of different dimensions to a weighted comparison range; and the weight coefficient is used to adjust the proportion of the model accuracy gain term and the resource constraint gain term in the reward.
[0091] The definition Value .
[0092] in, Indicates the current population state. Indicates an action, Indicates the current population state Next action The corresponding action value, Indicates the current population state State value, Indicates the current population state Next action The Q-value, relative to the advantage value of other actions, represents the long-term benefit that the deep Q-network estimates for choosing the corresponding subpopulation in the current state, and is used to evaluate the potential of each subpopulation as a source of breeding pool.
[0093] This allows the subpopulation selection process to simultaneously consider the structural characteristics of the current power AI model and the search direction of the compression strategy combination in a multi-objective optimization model. Compared to selecting subpopulations according to fixed probabilities or random methods, deep Q-networks can gradually adjust the evaluation of subpopulations based on the rewards obtained in each search process, making subsequent searches more focused on search directions that can improve inference accuracy or resource constraint indicators.
[0094] S7. Perform crossover and mutation processing on the compression strategy combinations in the breeding pool, and obtain an updated population through environmental selection; Offspring solutions are generated through crossover and mutation. The population is updated through environmental selection strategies, retaining individuals that meet the constraints and have better performance, and eliminating inferior solutions.
[0095] During the crossover process, two parent individuals are selected from the current population using either the sorting selection method or the Softmax probability selection method, and are designated as the first parent individual and the second parent individual. One or more crossover points are randomly selected, and the two parent individuals are combined using arithmetic crossover or heuristic crossover with compression strategies to generate offspring individuals.
[0096] In this context, the first parent individual and the second parent individual represent two candidate compression strategy combinations participating in the crossover operation; the crossover point represents the position in the two parent individuals used to exchange or combine strategy parameters; and the child individual represents the new candidate compression strategy combination formed after the crossover.
[0097] During the mutation process, Gaussian mutation is used to add Gaussian noise to the original values to change at least one gene in the offspring individuals, so as to generate a compression strategy combination that is different from that of the parent individuals.
[0098] The environment selection strategy is based on non-dominated ranking. (The parent population...) With offspring population A temporary population of size 2N is formed by merging the individuals. A fast non-dominated sort is performed on this temporary population, and the crowding distance of each individual in each non-dominated layer is calculated. The top N individuals are selected from the temporary population according to the non-dominated sort priority from high to low and the crowding distance from large to small to form a new generation population. This strategy ensures that while approaching the Pareto optimal frontier, it can maintain a uniform distribution of the solution set.
[0099] By cross-combining existing compression strategy combinations with their compression techniques and parameters, introducing new candidate compression strategy combinations through mutation, and retaining individuals with high priority on multiple evaluation indicators and relatively dispersed distribution in the target space through environmental selection, the search scope is expanded while preventing the population from prematurely concentrating in local search areas.
[0100] S8. Calculate the reward based on the updated population performance index, and update the deep Q network and experience pool based on the reward; The model accuracy gain term represents the change in inference accuracy of the child solution relative to the parent solution; the resource constraint gain term represents the change in computational cost, storage, latency, and power consumption of the child solution relative to the parent solution. The model accuracy gain term and the resource constraint gain term are normalized and then combined according to corresponding weight coefficients, which are determined by the fitness function.
[0101] Based on the updated population performance indicators Calculate rewards The value is updated to a depth-Q network, and the current population state, the selected subpopulation, the reward, and the updated population state are written to the experience pool. For the sake of reasoning accuracy, For computational load, For storage, For time delay, This refers to power consumption.
[0102] The performance metrics include inference accuracy, computational load, storage, latency, and power consumption. Inference accuracy represents the correctness of the model output, computational load represents the floating-point operation scale of one forward propagation, storage represents the memory usage required for model deployment and operation, latency represents the end-to-end inference time, and power consumption represents the computation and memory access power consumption during the inference process.
[0103] The impact of compression strategy combinations on model inference accuracy and terminal resource consumption is fed back to the deep Q-network. Therefore, in subsequent iterations, the deep Q-network can adjust its evaluation of subpopulations based on existing search results, avoiding the reliance on model accuracy or a single resource metric as the sole basis for subpopulation selection.
[0104] S9. Iterate through S6 to S8 until the termination condition is met, output the optimal compression strategy combination, and compress the power AI model according to the optimal compression strategy combination to obtain a lightweight model.
[0105] The termination conditions include reaching the maximum number of iterations, or the improvement in inference accuracy corresponding to the optimal compression strategy combination generated in P consecutive generations being less than a preset improvement threshold. P is a preset positive integer, set according to the search stability requirements. The improvement in inference accuracy is the positive increment of the inference accuracy corresponding to the current generation's optimal compression strategy combination relative to the inference accuracy corresponding to the previous generation's optimal compression strategy combination; when the current generation's inference accuracy is not higher than the previous generation's inference accuracy, the improvement in inference accuracy is 0.
[0106] When the termination condition is met, the optimal compression strategy combination that satisfies the constraints of the multi-objective optimization model is determined from the current population, and the power AI model is subjected to at least one of the following processes: weight pruning, quantization, sparse decomposition, and special architecture layer replacement, to obtain a lightweight model.
[0107] This enables the model compression process to ultimately output a lightweight model that can be directly used for deployment in smart power terminals, rather than just an abstract strategy search result, thus forming a complete closed loop from hardware resource awareness, business requirement modeling, compression strategy search to lightweight model generation.
[0108] Furthermore, a model adaptive compression method based on deep reinforcement learning can be employed. This method models the user's desired performance and hardware resource constraints as optimization objectives and constraints, formalizing the objectives and constraints into a unified optimization problem. Then, deep reinforcement learning is used to solve for hyperparameters and employ compression techniques to output the optimal compressed model. While the deep reinforcement learning-based model adaptive compression method has strong local fine-grained search capabilities, it also suffers from the limitation of easily getting trapped in local optima.
[0109] Please see Figure 2 In another embodiment of the present invention, an adaptive compression system for artificial intelligence models for power smart terminals is provided. This system can be used to implement the above-mentioned adaptive compression method for artificial intelligence models for power smart terminals. Specifically, the adaptive compression system for artificial intelligence models for power smart terminals includes an architecture parsing module, a resource quota mapping module, a constraint modeling module, and a model compression module.
[0110] Among them, the architecture parsing module is used to perform network architecture parsing on the trained power AI model to obtain the model structure features; The resource quota mapping module is used to map the underlying hardware parameters and system operation status of the power smart terminal to obtain a set of terminal resource quotas. The constraint modeling module is used to perform constraint modeling on the power business demand and the terminal resource quota set to obtain a multi-objective optimization model with the combination of compression strategies as decision variables. The model compression module is used to search for compression strategy combinations based on the multi-objective optimization model, the model structural features, and the terminal resource quota set to obtain the optimal compression strategy combination, and to compress the power AI model based on the optimal compression strategy combination to obtain a lightweight model.
[0111] The model compression module is used for: Direction vectors are generated in the target space of the multi-objective optimization model, and a population is constructed for the combination of compression strategies to obtain a set of direction vectors and an initial population. The compression strategy combinations in the initial population are partitioned by directional vector association to obtain multiple subpopulations; The current population state, which is composed of the structural features of the model and the direction vectors associated with each subpopulation, is evaluated by a deep Q-network to obtain the Q value of each subpopulation, and a breeding pool is obtained based on the Q value. Crossover and mutation are performed on the compression strategy combinations in the breeding pool to obtain offspring solutions, and environmental selection is performed on the parent population and the offspring solutions to obtain an updated population. The performance metrics of the updated population are used to calculate rewards, and the reward value is then used to update the deep Q-network and the experience pool. The deep Q-network evaluation, crossover and mutation, environment selection, and reward calculation are performed iteratively until the termination condition is met, thus obtaining the optimal compression strategy combination.
[0112] This invention presents an AI model adaptive compression system for power smart terminals. Through the collaborative work of four core modules—model acquisition, resource awareness, business requirement modeling, and model compression—it achieves automated adaptation and optimization from the original model to a lightweight terminal model. First, the architecture parsing module obtains the model from the cloud-based power business application or retrieves the original model to be compressed from the local AI model library, encapsulates it into a standardized data structure, and then passes it to subsequent modules. The resource quota mapping module collects hardware configuration information and runtime resource usage by calling the hardware abstraction layer interface of the terminal operating system. The constraint modeling module integrates the hardware parameters provided by the resource quota mapping module with the requirements of the business side, constructing a multi-objective optimization model by combining the inference accuracy threshold and latency threshold corresponding to the power business with the hardware resource quota. Finally, the model compression module, as the core execution unit, automatically searches for the optimal compression strategy combination using a multi-objective co-evolutionary algorithm and deep reinforcement learning. Under the premise of satisfying resource constraints, it quickly evaluates and selects the compression scheme with the best overall performance, ultimately outputting a lightweight model file, completing the adaptive compression of the model.
[0113] Please see Figure 3 and Figure 4This invention provides a smart power terminal device, which includes a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve corresponding method flows or corresponding functions. The processor described in this embodiment can be used for the operation of an adaptive compression method for artificial intelligence models in smart power terminals, including: The trained power AI model undergoes network architecture analysis to obtain model structural features; resource quota mapping is performed on the underlying hardware parameters of the power smart terminal to obtain a terminal resource quota set; constraint modeling is performed on power business requirements and the terminal resource quota set to obtain a multi-objective optimization model with compression strategy combinations as decision variables; direction vectors are generated in the objective space of the multi-objective optimization model, and a population is constructed for the compression strategy combinations to obtain a direction vector set and an initial population; the compression strategy combinations in the initial population are partitioned by direction vector association to obtain multiple sub-populations; the model structural features and each sub-population are then analyzed. The current population state, formed by the direction vectors of population association, is evaluated using a deep Q-network to obtain the Q-values of each subpopulation. A breeding pool is then selected based on these Q-values. Crossover and mutation are performed on the compression strategy combinations in the breeding pool to generate offspring solutions. An environment selection process is then applied to the parent population and the offspring solutions to obtain an updated population. Rewards are calculated for the performance indicators of the updated population to obtain reward values. These reward values are then used to update the deep Q-network and the experience pool. This process is repeated until a termination condition is met to obtain the optimal compression strategy combination. The power AI model is then compressed based on this optimal compression strategy combination to obtain a lightweight model.
[0114] Please see Figure 5The terminal device is a computer device. In this embodiment, the computer device 60 includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When executed by the processor 61, the computer program 63 implements the adaptive compression method for artificial intelligence models of power smart terminals described in this embodiment. To avoid repetition, these details are not elaborated here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the adaptive compression system for artificial intelligence models of power smart terminals described in this embodiment. To avoid repetition, these details are not elaborated here.
[0115] Computer device 60 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will understand that... Figure 5 This is merely an example of computer device 60 and does not constitute a limitation on computer device 60. It may include more or fewer components than shown, or combine certain components, or different components. For example, computer device may also include input / output devices, network access devices, buses, etc.
[0116] The processor 61 may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0117] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computer device 60.
[0118] Furthermore, the memory 62 may include both internal storage units of the computer device 60 and external storage devices. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.
[0119] Please see Figure 6 The terminal device is an electronic device 600, which is manifested in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.
[0120] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the method section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.
[0121] Storage unit 620 may include readable media in the form of volatile storage units, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include read-only memory (ROM) 6203.
[0122] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0123] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.
[0124] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem). This communication can be performed via input / output interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network, wide area network, and / or public network, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.
[0125] This invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the terminal device and extended storage media supported by the terminal device; it can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). More specific examples of the computer-readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical fiber, portable compact disk read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.
[0126] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency, etc., or any suitable combination thereof.
[0127] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0128] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the adaptive compression method for artificial intelligence models for power intelligent terminals in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor in the following steps: The trained power AI model undergoes network architecture analysis to obtain model structural features; resource quota mapping is performed on the underlying hardware parameters and system operation status of the power smart terminal to obtain a terminal resource quota set; constraint modeling is performed on power business requirements and the terminal resource quota set to obtain a multi-objective optimization model with compression strategy combinations as decision variables; direction vectors are generated in the objective space of the multi-objective optimization model, and a population is constructed for the compression strategy combinations to obtain a direction vector set and an initial population; the compression strategy combinations in the initial population are partitioned by direction vector association to obtain multiple sub-populations; the model structural features are then analyzed. The current population state, formed by the direction vectors associated with each subpopulation, is evaluated using a deep Q-network to obtain the Q-value of each subpopulation. A breeding pool is then selected based on these Q-values. Crossover and mutation are performed on the compression strategy combinations in the breeding pool to generate offspring solutions. An environment selection process is then applied to the parent population and the offspring solutions to obtain an updated population. Rewards are calculated for the performance indicators of the updated population to obtain reward values. These reward values are then used to update the deep Q-network and the experience pool. This process is repeated until a termination condition is met to obtain the optimal compression strategy combination. The power AI model is then compressed based on this optimal compression strategy combination to obtain a lightweight model.
[0129] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0130] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0131] To illustrate the function of the technical solution of this invention, the following explanation, in conjunction with the model compression process, focuses on hardware-software collaboration, compression strategy search, and model deployment adaptability.
[0132] I. Co-constraints between Model Performance and Terminal Resources Existing model compression methods typically focus on a single metric, such as model accuracy, number of parameters, or computational cost. Even if the compressed model meets the accuracy requirements, it may still be undeployable due to computational cost, storage usage, inference latency, or power consumption exceeding the limitations of terminal resources.
[0133] This invention acquires the network architecture parameters of the power AI model and the underlying hardware parameters and system operation status of the power smart terminal. Based on the underlying hardware parameters and system operation status, it sets computing resource quotas, memory quotas, and power consumption quotas, and sets inference accuracy thresholds and latency thresholds according to power business requirements. On this basis, using a combination of compression strategies as decision variables, it constructs a multi-objective optimization model that includes inference accuracy, computational load, storage, latency, and power consumption.
[0134] Compression strategy combinations are constrained by both model performance metrics and terminal resource boundaries during the search process. Only compression strategy combinations that meet constraints on inference accuracy, computational cost, storage, latency, and power consumption can be considered as candidate solutions for subsequent population updates. Therefore, the model compression process no longer only optimizes the model algorithm but also considers the model structure, terminal hardware resources, and power service requirements, ensuring that the generated lightweight model can run on the target power smart terminal.
[0135] II. Multi-directional search using a combination of compression strategies Weight pruning, quantization, sparse decomposition, and special architecture layer replacement can be used individually or in combination, each with different parameters. As the number of compression techniques and parameters increases, the number of compression strategy combinations also increases. When searching using a fixed process or a finite set of actions, feasible solutions are easily missed.
[0136] This invention generates direction vectors in a target space comprised of inference accuracy, computational cost, storage, latency, and power consumption. Based on the correlation between compression strategy combinations and these direction vectors, the initial population is divided into multiple subpopulations. Each subpopulation corresponds to a different target direction, enabling candidate compression strategy combinations to be searched around different performance trade-offs.
[0137] Through the above division, the combination of compression strategies will not be concentrated on a single evaluation metric. For example, one subpopulation can focus on inference accuracy, while another subpopulation can focus on computational cost, storage, latency, or power consumption, thereby retaining candidate schemes with different performance orientations and providing clear search units for subsequent deep Q-network selection of subpopulations.
[0138] III. Guiding Subpopulation Selection by Deep Q-Networks This invention uses the network architecture parameters of the currently optimized power AI model and the direction vector associated with the current compression strategy combination as input to the current population state into a deep Q-network. The deep Q-network calculates the Q value corresponding to each subpopulation and selects the subpopulations to form a breeding pool based on the Q value.
[0139] The network architecture parameters reflect the structural information of the power AI model to be compressed, such as weight parameters, activation value parameters, layer type, and number of layers. The direction vector reflects the search direction of the compression strategy combination in the target space. The deep Q-network evaluates different subpopulations based on the above information, so that the selection of subpopulations corresponds to the current model structure and search state, avoiding complete reliance on random methods or fixed rules to select breeding targets.
[0140] As iterations proceed, the deep Q-network updates its parameters based on the rewards corresponding to different subpopulations, thereby adjusting the subsequent evaluation results for the subpopulations. The search process can change the composition of the breeding pool based on the generated compression results, gradually updating the candidate solutions in a direction that meets the terminal resource constraints and business requirements.
[0141] IV. Synergistic Effects of Crossover, Variation, and Environmental Selection Crossing the compression strategy combinations in the breeding pool allows for the recombination of compression techniques and corresponding parameters from different parent individuals; mutation of the offspring individuals generated by the crossover can change some parameters or compression technique configurations, forming new compression strategy combinations.
[0142] After merging the parent and offspring populations, selection is performed using an environmental selection strategy based on non-dominated ordination and crowding distance. Non-dominated ordination is used to compare the priority relationships of each compression strategy combination on multiple evaluation metrics, while crowding distance is used to determine the distribution of each compression strategy combination in the target space.
[0143] By expanding candidate schemes through crossover and mutation, and then retaining individuals that meet the constraints and have different performance characteristics through environmental selection, the search scope can be expanded while avoiding the population from concentrating too early in a local area and maintaining the distribution of compression strategy combinations in the target space.
[0144] V. Reward-based search result feedback This invention calculates rewards based on the performance metrics of the updated population. The rewards consist of a model accuracy gain term and a resource constraint gain term. The model accuracy gain term reflects the change in inference accuracy of the child solution relative to the parent solution, while the resource constraint gain term reflects the changes in computational cost, storage, latency, and power consumption of the child solution.
[0145] The deep Q-network is updated based on the rewards, and the corresponding search information is stored in the experience pool. The results of the current iteration can be fed back into the subsequent subpopulation selection process. When a combination of compression strategies generated by a certain subpopulation improves the model accuracy or resource constraints, the corresponding search results will affect the Q-value of the subsequent output of the deep Q-network.
[0146] This feedback process enables deep Q-networks to adjust the selection of subpopulations using existing search results, rather than repeating the same selection method in each iteration, thereby reducing repeated searches for compression strategy combinations that do not meet the constraints.
[0147] VI. Adaptation to different smart power terminals Different smart power terminals differ in processor type, computing power, available memory, and allowable power consumption. A fixed compression scheme pre-set for a certain terminal is usually not directly applicable to terminals with other hardware configurations.
[0148] This invention addresses the acquisition of underlying hardware parameters and system operating status by current power smart terminals, and accordingly sets computing resource quotas, memory quotas, and power consumption quotas. It also sets inference accuracy thresholds and latency thresholds based on specific power services. For the same power AI model, when the target power smart terminal changes, a multi-objective optimization model can be reconstructed based on the new hardware parameters and service requirements, and a corresponding compression strategy combination can be searched.
[0149] Therefore, the same power AI model can obtain corresponding lightweight models for different power smart terminals, without the need to pre-fix a unique pruning ratio, quantization method or compression process for various terminals, thereby reducing the work of repeatedly adjusting manual parameters during cross-terminal deployment.
[0150] VII. Output of Model Compression Results After the termination condition is met, this invention outputs the optimal compression strategy combination that satisfies the constraints of the multi-objective optimization model, and performs corresponding weight pruning, quantization, sparse decomposition, or special architecture layer replacement on the power AI model according to the optimal compression strategy combination to obtain a lightweight model.
[0151] This process connects terminal hardware parameter acquisition, business requirement modeling, compression strategy combination search, and lightweight model generation, ultimately outputting a model file that can be deployed in power smart terminals, thus avoiding the model compression strategy remaining only in the algorithm search results without being implemented in the actual model.
[0152] In summary, this invention provides an adaptive compression method and system for artificial intelligence models of power smart terminals. By incorporating the network architecture parameters of the power AI model, the computing resource quotas, memory quotas, and power consumption quotas of the power smart terminal, as well as the inference accuracy thresholds and latency thresholds corresponding to power services, into a multi-objective optimization model, it achieves collaborative modeling of model structure, hardware resources, and service requirements. Subpopulations are divided using direction vectors, a deep Q-network is used to select the breeding pool, and the compression strategy combination is iteratively updated using crossover, mutation, and environmental selection. This allows the search process to adjust the search direction according to the model structure and resource constraints. Consequently, a compression strategy combination that meets the terminal resource conditions and service requirements can be output, resulting in a lightweight model suitable for deployment on different power smart terminals, reducing manual parameter tuning and repetitive adaptation work.
[0153] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0154] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0155] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0156] In the embodiments provided by this invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0157] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0158] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0159] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random-access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0160] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0161] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0162] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0163] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.
Claims
1. An adaptive compression method for artificial intelligence models of power intelligent terminals, characterized in that, When applied to smart power terminals, the following steps are included: S1. Perform network architecture analysis on the trained power AI model to obtain the model structure features; S2. Perform resource quota mapping on the underlying hardware parameters of the power smart terminal to obtain a terminal resource quota set; S3. Constraint modeling is performed on the power business demand and the set of terminal resource quotas to obtain a multi-objective optimization model with the combination of compression strategies as decision variables; S4. Generate direction vectors in the target space of the multi-objective optimization model, and construct a population of compression strategy combinations to obtain a set of direction vectors and an initial population. S5. Perform directional vector association partitioning on the compression strategy combinations in the initial population to obtain multiple subpopulations; S6. The current population state, which is composed of the structural features of the model and the direction vectors associated with each subpopulation, is evaluated by a deep Q-network to obtain the Q value of each subpopulation, and a breeding pool is obtained by filtering according to the Q value. S7. Perform crossover and mutation on the compression strategy combinations in the breeding pool to generate offspring solutions, and perform environmental selection on the parent population and the offspring solutions to obtain an updated population. S8. Calculate the reward for the performance index of the updated population to obtain the reward value, and update the deep Q network and experience pool according to the reward value. S9. Repeat S6 to S8 until the termination condition is met to obtain the optimal compression strategy combination, and compress the power AI model based on the optimal compression strategy combination to obtain a lightweight model.
2. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, The power AI model includes a new energy output prediction model, a load prediction model, an electricity consumption behavior analysis model, an intelligent inspection model, an online monitoring model, a fault analysis model, a distribution network topology identification model, or a regulation strategy generation model.
3. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, The model structure features include at least one of the following: weight parameters, activation value parameters, layer type, and number of layers in the power AI model.
4. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, In S2, the underlying hardware parameters are obtained by calling the hardware abstraction layer interface of the power smart terminal, and the computing resource quota, memory quota and power consumption quota are set according to the underlying hardware parameters and system operation status to obtain the terminal resource quota set.
5. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, In S3, the power service requirements include the inference accuracy threshold and latency threshold corresponding to the power service, and the terminal resource quota set includes computing resource quota, memory quota and power consumption quota. For any of the compression strategy combinations, the objective space of the multi-objective optimization model includes the inference accuracy, computational load, storage, latency, and power consumption of the power AI model after processing by the compression strategy combination. The constraints of the multi-objective optimization model include: the inference accuracy is not lower than the inference accuracy threshold, the computational load does not exceed the computational resource quota, the storage does not exceed the memory quota, the latency does not exceed the latency threshold, and the power consumption does not exceed the power consumption quota. The inference accuracy is: the ratio of the number of samples whose inferred values are equal to the true values to the total number of samples in the validation set when the power AI model after the compression strategy is combined and processed to infer the validation set samples. The computational load is determined based on the total number of floating-point arithmetic operations performed by the power AI model after processing by the compression strategy during a forward propagation process. The storage is determined based on the model parameters, intermediate activation values, and total memory required by the runtime environment of the power AI model after processing by the compression strategy. The delay is determined based on the sum of the computation time, data and instruction reading time, and scheduling time required for the power AI model after the compression strategy to process from receiving input data to outputting inference results; The power consumption is determined by the sum of the power consumption of the power AI model during the inference process and the power consumption of memory read / write during the compression strategy combined processing.
6. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, S5 include: Calculate the angle between each compression strategy combination in the initial population and each direction vector in the set of direction vectors; Each compression strategy combination is assigned to a subpopulation corresponding to the direction vector with the smallest included angle.
7. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, In S6, the state of the deep Q network includes the model structure features and the direction vector associated with the current compression strategy combination to be evaluated, and the action of the deep Q network includes selecting a subpopulation from multiple subpopulations; In S8, the reward value is determined based on the model accuracy gain term and the resource constraint gain term. The model accuracy gain term is used to represent the change in inference accuracy of the child solution relative to the parent solution, and the resource constraint gain term is used to represent the change in computation, storage, latency and power consumption of the child solution relative to the parent solution. The model accuracy gain term and the resource constraint gain term are normalized and then combined according to their corresponding weight coefficients, which are determined by the fitness function.
8. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, In S7: Two parent individuals are selected from the current population using either sorting selection or Softmax probabilistic selection. The compression strategies of the two parent individuals are then combined using arithmetic crossover or heuristic crossover to generate the offspring solution. Gaussian noise is added to the original values using Gaussian mutation, thereby changing at least one strategy parameter in the compression strategy combination corresponding to the sub-solution.
9. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, The environment selection is based on non-dominated sorting. In S7, environment selection is performed on the parent population and the offspring solutions to obtain an updated population, including: The parent population and the offspring population are merged to form a temporary population; The temporary population is subjected to fast non-dominated sorting, and the crowding distance of each individual in each non-dominated layer is calculated. According to the non-dominated sorting priority from high to low and the crowding distance from large to small, a preset number of individuals are selected from the temporary population to form the updated population.
10. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, The compression strategy combination includes at least one compression technique among weight pruning, quantization, sparse decomposition, and special architecture layer replacement.
11. The adaptive compression method for artificial intelligence models of power intelligent terminals according to claim 1, characterized in that, The termination conditions include reaching the maximum number of iterations, or the improvement in inference accuracy corresponding to the optimal compression strategy combination generated in consecutive P generations being less than a preset improvement threshold; where P is a preset positive integer, and the improvement in inference accuracy is the positive increment of the inference accuracy corresponding to the current generation's optimal compression strategy combination relative to the inference accuracy corresponding to the previous generation's optimal compression strategy combination.
12. An adaptive compression system for artificial intelligence models of power intelligent terminals, characterized in that, The system for performing the method according to any one of claims 1 to 11, the system comprising: The architecture parsing module is used to perform network architecture parsing on the trained power AI model to obtain the model's structural features. The resource quota mapping module is used to map the underlying hardware parameters and system operation status of the power smart terminal to obtain a set of terminal resource quotas. The constraint modeling module is used to perform constraint modeling on the power business demand and the terminal resource quota set to obtain a multi-objective optimization model with the combination of compression strategies as decision variables. The model compression module is used to search for compression strategy combinations based on the multi-objective optimization model, the model structural features, and the terminal resource quota set to obtain the optimal compression strategy combination, and to compress the power AI model based on the optimal compression strategy combination to obtain a lightweight model.
13. The adaptive compression system for artificial intelligence models of power intelligent terminals according to claim 12, characterized in that, The model compression module is used for: Direction vectors are generated in the target space of the multi-objective optimization model, and a population is constructed for the combination of compression strategies to obtain a set of direction vectors and an initial population. The compression strategy combinations in the initial population are partitioned by directional vector association to obtain multiple subpopulations; The current population state, which is composed of the structural features of the model and the direction vectors associated with each subpopulation, is evaluated by a deep Q-network to obtain the Q value of each subpopulation, and a breeding pool is obtained based on the Q value. Crossover and mutation are performed on the compression strategy combinations in the breeding pool to obtain offspring solutions, and environmental selection is performed on the parent population and the offspring solutions to obtain an updated population. The performance metrics of the updated population are used to calculate rewards, and the reward value is then used to update the deep Q-network and the experience pool. The deep Q-network evaluation, crossover and mutation, environment selection, and reward calculation are performed iteratively until the termination condition is met, thus obtaining the optimal compression strategy combination.
14. A smart power terminal, characterized in that, Including processor and memory; The memory stores a computer program, which, when executed by the processor, causes the processor to perform the method described in any one of claims 1 to 11.
15. The power intelligent terminal according to claim 14, characterized in that, The processor includes at least one of a central processing unit, a graphics processing unit, a neural network processing unit, or a field-programmable gate array; The memory also stores the operating system of the power smart terminal. The operating system includes a kernel layer, a system service layer, and an application layer. When the computer program is executed by the processor, it obtains the underlying hardware parameters of the power smart terminal through the interface of the operating system and deploys the compressed lightweight model on the system service layer for edge computing services.
16. A chip, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1 to 11.