Industrial ai model compression method based on hybrid precision quantization and adaptive sparsification

CN122840130APending Publication Date: 2026-09-29GUANGZHOU CHUANGBO MECH & ELECTRICAL EQUIP INSTALLATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611258333.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-19
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0002]随着工业4.0和智能制造的快速发展,深度学习模型已广泛应用于工业缺陷检测、设备预测性维护、机器人路径规划等关键生产环节,工业生产线往往面临多种不同的使用场景切换,传统的模型压缩方法如量化、剪枝、知识蒸馏等虽能在一定程度上减小模型体积,但往往采用“一刀切”的全局压缩策略,未能充分考虑工业场景多样性,从而影响工业AI模型对使用场景的适配能力

Benefits of technology

[0041]通过质点在网络层驻点之间的迁移,实现位宽资源的流动分配。每个卷积通道根据当前使用场景的重要性得分获得对应数量的质点即对应位宽,重要度高的通道分配更多质点更高位宽,重要度低的通道分配较少或零质点加以剪枝。该机制实现了从层级别粗粒度分配到通道级别细粒度分配的跨越,提高位宽资源的利用效率。同时,通过将工业AI模型的位宽等分成多份并与质点绑定,使CPU、GPU、NPU等计算单元保持在高效工作区间,单位时间内处理更多工业数据样本,从而提高工业AI模型对使用场景的适配能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840130A_ABST
    Figure CN122840130A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of model compression, in particular to an industrial AI model compression method based on mixed precision quantization and adaptive sparsification, which comprises the following steps: obtaining a plurality of AI sub-models corresponding to a plurality of different use scenarios respectively, extracting model characteristics corresponding to each AI sub-model, fusing the plurality of AI sub-models based on the model characteristics to construct an industrial AI model; configuring a compression assistant for the industrial AI model, wherein the compression assistant comprises a plurality of mass points and corresponding moving channels; obtaining a current use scenario of the industrial AI model in real time, and compressing the industrial AI model according to the compression assistant to obtain a target AI model corresponding to the current use scenario, so that the bit width of the industrial AI model can be divided into multiple parts and bound to the mass points, the computing unit can be kept in an efficient working interval, more industrial data samples can be processed in a unit of time, and the adaptation capability of the industrial AI model to the use scenario is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of model compression technology, specifically to an industrial AI model compression method based on hybrid precision quantization and adaptive sparsity. Background Technology

[0002] With the rapid development of Industry 4.0 and intelligent manufacturing, deep learning models have been widely used in key production processes such as industrial defect detection, predictive maintenance of equipment, and robot path planning. Industrial production lines often face a variety of different usage scenarios. Although traditional model compression methods such as quantization, pruning, and knowledge distillation can reduce the model size to some extent, they often adopt a "one-size-fits-all" global compression strategy, which fails to fully consider the diversity of industrial scenarios, thus affecting the adaptability of industrial AI models to usage scenarios.

[0003] To address these issues, we propose an industrial AI model compression method based on hybrid precision quantization and adaptive sparsity. Summary of the Invention

[0004] The purpose of this invention is to provide an industrial AI model compression method based on hybrid precision quantization and adaptive sparsity to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an industrial AI model compression method based on hybrid precision quantization and adaptive sparsity, the method comprising the following steps:

[0006] Obtain multiple AI sub-models corresponding to different use scenarios, extract the model features of each AI sub-model, and merge the multiple AI sub-models based on the model features to construct an industrial AI model.

[0007] Configure a compression assistant for the industrial AI model, which includes multiple mass points and corresponding movement channels;

[0008] The system acquires the current usage scenario of the industrial AI model in real time, and compresses the industrial AI model using a compression assistant to obtain the target AI model corresponding to the current usage scenario.

[0009] Preferably, the steps of acquiring multiple AI sub-models corresponding to multiple different usage scenarios, extracting model features corresponding to each AI sub-model, and fusing multiple AI sub-models based on model features to construct an industrial AI model include:

[0010] Multiple different usage scenarios are obtained, and the industrial AI model is trained separately according to the different usage scenarios to obtain multiple AI sub-models;

[0011] Obtain the common features of each AI sub-model, including the same network layer and the same convolutional channels on the network layer, and share the common features corresponding to multiple use cases to build the basic network layer.

[0012] Obtain the individual features corresponding to each AI sub-model. The individual features include the independent network layers and convolutional channels corresponding to each AI sub-model.

[0013] The individual features corresponding to multiple AI sub-models are fused into the basic network layer to obtain an industrial AI model.

[0014] Preferably, the step of configuring a compression assistant for the industrial AI model, wherein the compression assistant includes multiple mass points and corresponding movement channels, includes:

[0015] Based on multiple AI sub-models, the multiple convolutional channels on multiple network layers in the industrial AI model are divided into convolutional channel clusters, where each AI sub-model corresponds to at least one convolutional channel cluster.

[0016] The multiple convolutional channel clusters are labeled respectively, wherein the convolutional channel clusters corresponding to different AI sub-models have different labels;

[0017] Obtain multiple network layers from the industrial AI model and assign multiple mass points to each network layer;

[0018] Establish connection channels between multiple network layers, and select the corresponding connection channel as the movement channel of the mass point according to the usage scenario.

[0019] Preferably, the steps of obtaining multiple network layers in the industrial AI model, assigning multiple mass points to each network layer, and establishing connection channels between the multiple network layers include:

[0020] For each convolutional channel of multiple network layers, multiple stationary points are set, and multiple particles with the same number as the stationary points of the first network layer are set and located on the corresponding stationary points;

[0021] Multiple stations located on the same network layer are used as information layers. Connection lines are established between stations on two adjacent network layers, and multiple network layers are connected through multiple connection lines to obtain the connection channel between network layers.

[0022] Preferably, the step of acquiring the current usage scenario of the industrial AI model in real time and compressing the industrial AI model using a compression assistant to obtain the target AI model corresponding to the current usage scenario includes:

[0023] Establish mapping relationships between multiple use cases and multiple AI sub-models, and store them in the scene-model mapping table of the compression assistant;

[0024] Real-time acquisition of the current application scenario of industrial AI models, and acquisition of AI sub-models corresponding to the current application scenario based on mapping relationships;

[0025] By labeling the convolutional channel clusters corresponding to the AI ​​sub-models in the industrial AI model with particles, and pruning the convolutional channels corresponding to unlabeled stationary points, the compressed target AI model is obtained.

[0026] Preferably, the step of labeling the convolutional channel clusters corresponding to the AI ​​sub-model in the industrial AI model using mass points includes:

[0027] Obtain the position information of the AI ​​sub-model in the industrial AI model, and bind the position information to each mass point;

[0028] The information layer is obtained by connecting the stationary points of the convolution channels of the first network layer corresponding to the AI ​​sub-model in the industrial AI model, and multiple particles are laid on multiple stationary points in the information layer.

[0029] Multiple points determine multiple convolution channels for corresponding position information in the next network layer, and establish temporary connection lines between the current stationary point and the corresponding stationary point in the next network layer as temporary connection channels. Multiple stationary points located in the same network layer are interconnected.

[0030] Multiple points are moved from their current stationary point to the next network layer via temporary connection channels, and the importance of each convolution channel is obtained.

[0031] The number of particles is allocated according to the importance of the convolution channels; the stationary points where multiple particles stay are marked, and the convolution channels corresponding to the marked stationary points are combined to obtain convolution channel clusters.

[0032] Preferably, the step of pruning the convolutional channels corresponding to unlabeled stationary points to obtain the compressed target AI model includes:

[0033] In industrial AI models, the stationary points corresponding to convolutional channels outside the convolutional channel clusters are marked as unmarked stationary points.

[0034] Marked stations adjacent to unmarked stations are used as edge stations of the target AI model. Multiple edge stations are connected to obtain the contour line of the target AI model in the industrial AI model. Based on the contour line, multiple unmarked stations are pruned.

[0035] The pruned industrial AI model is used as the compressed target AI model.

[0036] Preferably, the step of pruning multiple unmarked stationary points based on the contour line includes:

[0037] For each station, a port pair is set, the port pair including a first port and a second port;

[0038] Divide the port pair into two ports at the outline and temporarily disconnect the connection during the pruning operation;

[0039] When it is necessary to restore the pruned convolutional channels, the pruned convolutional channels are reconnected with the corresponding retained convolutional channels according to the port pairs to generate the restored industrial AI model.

[0040] Compared with the prior art, the beneficial effects of the present invention are:

[0041] By migrating particles between network layer stations, bit width resources are dynamically allocated. Each convolutional channel receives a corresponding number of particles (i.e., a corresponding bit width) based on its importance score for the current use case. Channels with higher importance are allocated more particles and higher bit widths, while channels with lower importance are allocated fewer particles or zero particles (pruning). This mechanism achieves a leap from coarse-grained allocation at the layer level to fine-grained allocation at the channel level, improving the utilization efficiency of bit width resources. Simultaneously, by equally dividing the bit width of the industrial AI model into multiple parts and binding them to particles, computing units such as CPUs, GPUs, and NPUs remain in their efficient operating range, processing more industrial data samples per unit time, thereby improving the adaptability of the industrial AI model to different use cases. Attached Figure Description

[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] For examples, please refer to Figure 1 This invention provides a technical solution for industrial AI model compression based on mixed-precision quantization and adaptive sparsity: the industrial AI model compression method based on mixed-precision quantization and adaptive sparsity includes the following steps:

[0046] S1: Obtain multiple AI sub-models corresponding to different use scenarios, extract the model features corresponding to each AI sub-model, and merge multiple AI sub-models based on the model features to construct an industrial AI model;

[0047] The steps for obtaining multiple AI sub-models corresponding to different usage scenarios, extracting model features from each AI sub-model, and fusing these features to construct an industrial AI model include: obtaining multiple different usage scenarios; training the industrial AI model separately for each usage scenario to obtain multiple AI sub-models; obtaining common features from each AI sub-model, including the same network layers and the same convolutional channels within those layers, and sharing these common features across multiple usage scenarios to construct a base network layer; obtaining individual features from each AI sub-model, including the independent network layers and convolutional channels for each AI sub-model; and fusing these individual features from multiple AI sub-models into the base network layer to obtain the industrial AI model.

[0048] Specifically, differences in data distribution, defect types, and deployed hardware in different scenarios will change the "importance score" (degree of importance) of each network layer. The same layer may be a key feature extractor in scenario A, but may become redundant in scenario B.

[0049] For example, in scenario A: Suppose a general pre-trained model for metal surface defect detection contains 5 convolutional blocks (Block 1~Block 5). The data characteristics are fast-moving parts with motion blur, mainly large-area scratches (coarse-grained features), large defect area, high contrast, and easy detection. The sensitivity analysis results (calculated based on actual collected data) are: Block 1 has an extremely high importance score because it extracts basic edge texture features and relies more on low-level features under motion blur. Therefore, the corresponding compression strategy is high precision, which requires allocating more points to this layer. Since the number of points is tied to the bit width, more points mean more bit width. Blocks 2~3 have medium importance scores because they have medium semantic features with some redundancy. Therefore, the corresponding points can be appropriately reduced, thereby reducing the corresponding bit width, which is medium precision. Blocks 4~5 have very low importance scores because the high-level semantic feature map has low resolution. Large defect detection no longer requires fine semantics. Therefore, some local content can be pruned and fewer points can be allocated. High-speed production lines need to protect low-level channels and compress high-level channels.

[0050] Scenario B: Precision parts inspection (small defects, high precision), hardware constraints: local industrial PC (i7 + Quadro GPU), sufficient computing power, high precision is required without compromise; data characteristics: static high-definition imaging of parts; detection of small defects such as microcracks (<0.1mm) and pores; low contrast between defects and background; sensitivity analysis results show that the importance score of Block 1 is medium because the basic feature discrimination is not high and all parts have similar edges, so appropriate pruning, such as 20%, can be performed; the importance scores of Blocks 2-3 are extremely high because the low-to-medium level semantics can best distinguish small texture anomalies, which are key feature layers, so no pruning is required, and high precision is achieved (representing a large bit width allocation); the importance scores of Blocks 4-5 are extremely high because a large receptive field is needed to compare local and global features, small defects rely on high-level semantic localization, so no pruning is required, and a high bit width allocation is needed. Precision inspection requires protection of the middle and high-level channels, so compression is limited;

[0051] Each network layer contains multiple channels, and importance can be determined at the channel level.

[0052] For example, the 64th channel of Block 3 vs. the 128th channel:

[0053] In scenario A (high-speed production line):

[0054] Channel 64: Learns "short edges in the 45° direction" - crucial for scratch detection under motion blur → Retain;

[0055] Channel 128: Learns "dense dot texture" - it becomes blurred into a mass under high-speed imaging and loses its distinguishability → pruning;

[0056] In scenario C (precision detection):

[0057] Channel 64: Widespread presence at 45° edges, no special contribution to microcrack detection → can be pruned;

[0058] Channel 128: Dotted texture can highlight small defects such as sand holes and pores → Keep;

[0059] In summary, different layers and channels are compressed according to different scenarios. By analyzing the sensitivity and importance of convolutional channels, a corresponding compression strategy is customized for each use case, thereby making the resulting target AI model more suitable for the corresponding application scenario.

[0060] S2: Configure a compression assistant for the industrial AI model, where the compression assistant includes multiple mass points and corresponding movement channels;

[0061] The steps of configuring a compression assistant for an industrial AI model, wherein the compression assistant includes multiple mass points and corresponding movement channels, include: dividing multiple convolutional channels on multiple network layers in the industrial AI model into convolutional channel clusters according to multiple AI sub-models, wherein each AI sub-model corresponds to at least one convolutional channel cluster; labeling the multiple convolutional channel clusters respectively, wherein the convolutional channel clusters corresponding to different AI sub-models have different labels; obtaining multiple network layers in the industrial AI model and setting multiple mass points for each network layer; establishing connection channels between multiple network layers, and selecting the corresponding connection channels as the movement channels for the mass points according to the usage scenario;

[0062] The steps for obtaining multiple network layers in an industrial AI model and setting multiple mass points for each network layer to establish connection channels between multiple network layers include: setting multiple stationary points for each convolutional channel of multiple network layers, setting multiple mass points with the same number of stationary points as the first network layer, and placing them on the corresponding stationary points; taking multiple stationary points located on the same network layer as information layers, establishing connection lines between stationary points on two adjacent network layers, and connecting multiple network layers through multiple connection lines to obtain connection channels between network layers;

[0063] Specifically, a particle is a virtual computing unit carrying bit-width resources, used to migrate between network layers in an industrial AI model. It dynamically allocates quantized bit-width resources based on the importance assessment of convolutional channels in the current usage scenario. Each particle is bound to a unit bit-width (e.g., 1 bit or a group of bits). The number of particles is positively correlated with the importance of the convolutional channels: the higher the importance, the more particles are allocated, and the corresponding bit-width is higher. Particles move between resident points in the network layers, forming a fluid allocation of bit-width resources. Each particle is a virtual machine, carrying a computing quota and a memory bandwidth quota. A unit bit-width is the smallest precision unit supported by the hardware instruction set. Multiple particles form a resource scheduling cluster used to manage the total amount of allocable computing resources. A stationary point is a virtual anchorage location set at the input and output ports of each convolutional channel in an industrial AI model. It serves to hold particles and allocate bit width resources. Each convolutional channel corresponds to one input stationary point and one output stationary point. Each stationary point has a unique identifier (ID) used to locate the channel's position in the network topology. When a particle rests on a stationary point, it means that the corresponding channel receives a corresponding amount of bit width resources. A stationary point is a computing node or memory address pointer, and in virtual machines, it is bound to a blade-feature physical core or memory partition. An information layer is a collection of all stationsary points located on the same network layer. It is used to summarize and manage the particle distribution status and bit width configuration of all convolutional channels in that network layer. Multiple stationsary points within an information layer can be interconnected to achieve the reallocation of particle resources within that layer. The information layer serves as the basic unit for compression assistants to execute layer-level decisions; it is also a switch aggregation node that manages the resources of a group of physical servers. A connection channel is a fixed topology connection between stationsary points in two adjacent information layers, used to define the standard movement path of particles between network layers. Connection channels are pre-established during model initialization, reflecting the connection relationships of the original network structure, such as physical network links. Connection channels are static and preset, and do not change with the usage scenario. Temporary connection channels are temporary connections dynamically established by the compression assistant according to different usage scenarios to guide the migration of points between irregular paths. They are created and destroyed in real time by the compression assistant and are scenario-specific; they allow points to skip unimportant layers or directly enter critical layers, realizing shortcut allocation of computing resources (the load balancer directly directs traffic to critical service nodes). Temporary connection channels are also virtual machine dynamic migration channels.

[0064] S3: Real-time acquisition of the current usage scenario of the industrial AI model, and compression of the industrial AI model by the compression assistant to obtain the target AI model corresponding to the current usage scenario;

[0065] The steps for obtaining the target AI model corresponding to the current use scenario of the industrial AI model in real time and compressing the industrial AI model according to the compression assistant include: establishing a mapping relationship between multiple use scenarios and multiple AI sub-models and storing it in the scenario-model mapping table of the compression assistant; obtaining the current use scenario of the industrial AI model in real time and obtaining the AI ​​sub-model corresponding to the current use scenario based on the mapping relationship; marking the convolution channel cliques corresponding to the AI ​​sub-models in the industrial AI model by using mass points, and pruning the convolution channels corresponding to unmarked stationary points to obtain the compressed target AI model;

[0066] Establish mapping relationships between multiple use cases and multiple AI sub-models, and store them in the scene-model mapping table of the compression assistant. Specifically, the mapping relationships are learned automatically from historical scene data using meta-learning or transfer learning methods to study the correlation patterns between scenes and sub-models. The mapping relationships include one or more of the following types:

[0067] One-to-one mapping: One use case uniquely corresponds to one AI sub-model;

[0068] One-to-many mapping: One use case corresponds to multiple AI sub-models, and the optimal sub-model is dynamically selected based on real-time performance metrics;

[0069] Many-to-one mapping: Multiple use cases share the same AI sub-model, which is suitable for situations where the similarity between scenarios exceeds a preset threshold;

[0070] Hierarchical mapping: Use cases are classified by level, with parent scenarios corresponding to coarse-grained sub-models and sub-scenarios corresponding to fine-grained adjusted sub-model variants;

[0071] The scene-model mapping table includes fields such as scene identifier (a unique ID or code that identifies a use case), scene feature vector (a multi-dimensional feature description used for scene matching), corresponding sub-model ID (a unique identifier pointing to the corresponding AI sub-model), and default compression configuration (pre-calculated recommended compression parameters for this scene).

[0072] The steps of marking convolutional channel clusters corresponding to AI sub-models in an industrial AI model using particles include: obtaining the position information of the AI ​​sub-model in the industrial AI model and binding the position information to each particle; connecting the stationary points of the convolutional channels of the first network layer corresponding to the AI ​​sub-model in the industrial AI model to obtain an information layer, and placing multiple particles on multiple stationary points in the information layer; determining multiple convolutional channels corresponding to the position information of the next network layer for multiple particles, and establishing temporary connection lines between the stationary point where the current particle is located and the stationary point of the corresponding position information in the next network layer as temporary connection channels; connecting multiple stationary points located in the same network layer to each other; moving multiple particles from their current stationary point to the next network layer along the temporary connection channels, and obtaining the importance of each convolutional channel, allocating the corresponding number of particles according to the importance of the convolutional channels; marking the stationary points where multiple particles reside, and combining the convolutional channels corresponding to the marked stationary points to obtain convolutional channel clusters;

[0073] The steps for pruning the convolutional channels corresponding to unlabeled stationary points to obtain the compressed target AI model include: labeling the stationary points corresponding to convolutional channels outside the convolutional channel clusters in the industrial AI model as unlabeled stationary points; using the labeled stationary points adjacent to the unlabeled stationary points as edge stationary points of the target AI model; connecting multiple edge stationary points to obtain the contour line of the target AI model in the industrial AI model; pruning multiple unlabeled stationary points based on the contour line; and using the pruned industrial AI model as the compressed target AI model.

[0074] The steps of pruning multiple unmarked stationary points based on the contour line include: setting a port pair for each stationary point, the port pair including a first port and a second port; dividing the port pair into two ports at the contour line and temporarily disconnecting the connection during the pruning operation; when it is necessary to restore the pruned convolutional channels, according to the port pair, reconnecting the multiple pruned convolutional channels with the corresponding retained convolutional channels to generate the restored industrial AI model;

[0075] It should be noted that the first port is located on the side of the pruned convolution channel, storing a copy of the weight parameters and topology connection information of the convolution channel; the second port is located on the side of the retained convolution channel, storing the connection interface identifier and restoration anchor point information; the ports remain independent after being disconnected by pruning, and are reconnected through port matching when restoration is needed;

[0076] The temporary disconnection operation includes: maintaining the logical connection between the port pairs, but prohibiting the actual flow of data in the forward propagation; freezing the weight parameters of the disconnected convolutional channels and storing them in the compression assistant's cache, not participating in model inference calculation; and reconnecting the two ports of the port pair when the current usage scenario changes and the importance assessment of the pruned channels in the new scenario exceeds a preset threshold.

[0077] The recovery operation includes: locating the position of the pruned channel in the original network topology based on the connection interface identifier in the port pair; unfreezing the stored channel weight parameters; and reconnecting the channel to the forward propagation computation graph to form the recovered industrial AI model.

[0078] A port pair is a pair of virtual interfaces set at the outline, including a first port on the pruned side and a second port on the preserved side, used to enable recoverable disconnection and reconnection of the pruning channel. The first port is the backup storage volume snapshot, and the second port is the service discovery endpoint. When the two ports in the port pair are connected, it is in normal inference state; when the two ports are disconnected, it is in pruning state, the logical relationship is preserved but data does not flow; during recovery, a reconnection operation is performed. The port pair is the connection between the microservice gateway and the backend service instance; pruning disconnection is the gateway removing unhealthy / low-priority service instances (keeping the routing entries frozen); recovery reconnection is the gateway redirecting traffic to the recovered instance. The contour line is a virtual boundary connecting multiple edge stations, used to define the core computing area that needs to be retained and the peripheral area that can be pruned in the current use case. The contour line is the physical security boundary of the data center or the gateway line of the network firewall. The convolutional channel cluster is the smallest indivisible functional unit formed by the aggregation of multiple highly related and cooperating convolutional channels, corresponding to an AI sub-model in the current use case, and belongs to the available area. The compression assistant is a management unit configured on the industrial AI model side, containing multiple particles and corresponding movement channels, used to perceive the use case in real time and perform compression operations such as bit width allocation, particle scheduling, and pruning control. Through the migration of particles between network layer stations, the flow allocation of bit width resources is realized. Each convolutional channel obtains a corresponding number of particles, i.e., the corresponding bit width, according to the importance score of the current use case. Channels with higher importance are allocated more particles and higher bit widths, while channels with lower importance are allocated fewer or zero particles and are pruned. This mechanism realizes the leap from coarse-grained allocation at the layer level to fine-grained allocation at the channel level, improving the utilization efficiency of bit width resources. Meanwhile, by dividing the bit width of the industrial AI model into multiple equal parts and binding them to mass points, the computing units such as CPU, GPU, and NPU are kept in a high-efficiency working range, processing more industrial data samples per unit time, thereby improving the adaptability of the industrial AI model to the application scenario.

[0079] Specifically, multiple AI sub-models are first trained based on the application scenarios of the industrial AI model. Then, these sub-models are fused together based on their shared and differing features to obtain the industrial AI model. The industrial AI model contains multiple network layers, each with multiple convolutional channels. Each network layer sets a corresponding stationary point for each convolutional channel, serving as the dwell point for the particles in the AI ​​sub-models corresponding to different application scenarios. When distributing the particles on the first network layer, since the importance of each convolutional channel varies in the current application scenario, different numbers of particles need to be distributed. Higher importance channels receive more particles, and lower importance channels receive fewer particles. The lower the assigned mass point, the less necessary it is to place mass points at the corresponding stationary points when they are unrelated to the current use scenario. When determining the target AI model corresponding to the current use scenario from the industrial AI model, the position information of the AI ​​sub-model corresponding to the current use scenario in the target AI model is first obtained. This position information refers to the position information of each network layer and convolutional channel that the current use scenario needs to use in the industrial AI model. The position information of the convolutional channels is determined through the identity identifier of the stationary point. After obtaining the position information, different numbers of mass points are allocated according to the importance of each convolutional channel in the first network layer. After the analysis of the first network layer is completed, the maximum number of mass points is obtained. Multiple stationary points are established, and temporary connection channels are created between these stationary points and any stationary point in the second network layer. Here, any stationary point is a stationary point located in the position information of the target AI model. Multiple particles from the previous network layer are transferred to the second network layer along the temporary connection channels. Simultaneously, the positions of the multiple particles are allocated according to the importance of the convolutional channels corresponding to the multiple stationary points in the second network layer, moving the particles to their corresponding stationary points. The number of particles residing at a stationary point is ensured to be consistent with the importance of the convolutional channel. The bit width of the industrial AI model is divided into multiple equal parts, each bit width corresponding to one particle, and these parts are bound together. Different levels of importance stationary points correspond to different... The number of points can be allocated to different bit widths. Reducing the bit width can linearly reduce storage and bandwidth. Bit width is the number of bits occupied by each value. Reducing the bit width can convert high-precision operations into multiple low-precision operations, thereby improving computing power on hardware that supports low bit width. Mixed-precision quantization of industrial AI models focuses the bit width on important channels and network layers, reducing the waste of bit width for redundant features and avoiding bit width idleness or overload. This keeps computing units such as CPU / GPU / NPU in their efficient working range, processing more industrial data samples per unit time. It can ensure the rationality of bit width allocation in industrial AI models, thereby improving the adaptability of industrial AI models to usage scenarios.

[0080] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0081] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An industrial AI model compression method based on hybrid precision quantization and adaptive sparsity, characterized in that, Includes the following steps: Obtain multiple AI sub-models corresponding to different use scenarios, extract the model features of each AI sub-model, and merge the multiple AI sub-models based on the model features to construct an industrial AI model. Configure a compression assistant for the industrial AI model, which includes multiple mass points and corresponding movement channels; The system acquires the current usage scenario of the industrial AI model in real time, and compresses the industrial AI model using a compression assistant to obtain the target AI model corresponding to the current usage scenario.

2. The industrial AI model compression method based on hybrid precision quantization and adaptive sparsity according to claim 1, characterized in that: The steps of acquiring multiple AI sub-models corresponding to different application scenarios, extracting model features corresponding to each AI sub-model, and fusing multiple AI sub-models based on model features to construct an industrial AI model include: Multiple different usage scenarios are obtained, and the industrial AI model is trained separately according to the different usage scenarios to obtain multiple AI sub-models; Obtain the common features of each AI sub-model, including the same network layer and the same convolutional channels on the network layer, and share the common features corresponding to multiple use cases to build the basic network layer. Obtain the individual features corresponding to each AI sub-model. The individual features include the independent network layers and convolutional channels corresponding to each AI sub-model. The individual features corresponding to multiple AI sub-models are fused into the basic network layer to obtain an industrial AI model.

3. The industrial AI model compression method based on hybrid precision quantization and adaptive sparsity according to claim 1, characterized in that: The step of configuring a compression assistant for the industrial AI model, wherein the compression assistant includes multiple mass points and corresponding movement channels, includes: Based on multiple AI sub-models, the multiple convolutional channels on multiple network layers in the industrial AI model are divided into convolutional channel clusters, where each AI sub-model corresponds to at least one convolutional channel cluster. The multiple convolutional channel clusters are labeled respectively, wherein the convolutional channel clusters corresponding to different AI sub-models have different labels; Obtain multiple network layers from the industrial AI model and assign multiple mass points to each network layer; Establish connection channels between multiple network layers, and select the corresponding connection channel as the movement channel of the mass point according to the usage scenario.

4. The industrial AI model compression method based on hybrid precision quantization and adaptive sparsity according to claim 3, characterized in that: Obtain multiple network layers from the industrial AI model and assign multiple mass points to each network layer; The steps to establish a connection channel between multiple network layers include: For each convolutional channel of multiple network layers, multiple stationary points are set, and multiple particles with the same number as the stationary points of the first network layer are set and located on the corresponding stationary points; Multiple stations located on the same network layer are used as information layers. Connection lines are established between stations on two adjacent network layers, and multiple network layers are connected through multiple connection lines to obtain the connection channel between network layers.

5. The industrial AI model compression method based on hybrid precision quantization and adaptive sparsity according to claim 1, characterized in that: The steps of acquiring the current usage scenario of the industrial AI model in real time and compressing the industrial AI model using a compression assistant to obtain the target AI model corresponding to the current usage scenario include: Establish mapping relationships between multiple use cases and multiple AI sub-models, and store them in the scene-model mapping table of the compression assistant; Real-time acquisition of the current application scenario of industrial AI models, and acquisition of AI sub-models corresponding to the current application scenario based on mapping relationships; By labeling the convolutional channel clusters corresponding to the AI ​​sub-models in the industrial AI model with particles, and pruning the convolutional channels corresponding to unlabeled stationary points, the compressed target AI model is obtained.

6. The industrial AI model compression method based on hybrid precision quantization and adaptive sparsity according to claim 5, characterized in that: The step of labeling the convolutional channel clusters corresponding to the AI ​​sub-model in the industrial AI model using mass points includes: Obtain the position information of the AI ​​sub-model in the industrial AI model, and bind the position information to each mass point; The information layer is obtained by connecting the stationary points of the convolution channels of the first network layer corresponding to the AI ​​sub-model in the industrial AI model, and multiple particles are laid on multiple stationary points in the information layer. Multiple points determine multiple convolution channels for corresponding position information in the next network layer, and establish temporary connection lines between the current stationary point and the corresponding stationary point in the next network layer as temporary connection channels. Multiple stationary points located in the same network layer are interconnected. Multiple points are moved from their current stationary point to the next network layer via temporary connection channels, and the importance of each convolution channel is obtained. The number of particles is allocated according to the importance of the convolution channels; the stationary points where multiple particles stay are marked, and the convolution channels corresponding to the marked stationary points are combined to obtain convolution channel clusters.

7. The industrial AI model compression method based on hybrid precision quantization and adaptive sparsity according to claim 5, characterized in that: The step of pruning the convolutional channels corresponding to unlabeled stationary points to obtain the compressed target AI model includes: In industrial AI models, the stationary points corresponding to convolutional channels outside the convolutional channel clusters are marked as unmarked stationary points. Marked stations adjacent to unmarked stations are used as edge stations of the target AI model. Multiple edge stations are connected to obtain the contour line of the target AI model in the industrial AI model. Based on the contour line, multiple unmarked stations are pruned. The pruned industrial AI model is used as the compressed target AI model.

8. The industrial AI model compression method based on hybrid precision quantization and adaptive sparsity according to claim 7, characterized in that: The step of pruning multiple unmarked stationary points based on the contour line includes: For each station, a port pair is set, the port pair including a first port and a second port; Divide the port pair into two ports at the outline and temporarily disconnect the connection during the pruning operation; When it is necessary to restore the pruned convolutional channels, the pruned convolutional channels are reconnected with the corresponding retained convolutional channels according to the port pairs to generate the restored industrial AI model.