Model optimization method, communication equipment and storage medium
By employing model optimization methods and utilizing model pruning, quantization, and distillation techniques, the problem of deploying deep learning models on resource-constrained devices was solved, achieving efficient and adaptable model deployment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2025-03-17
- Publication Date
- 2026-05-01
AI Technical Summary
Deep learning models are difficult to deploy on resource-constrained devices, mainly due to high storage requirements and high computational complexity, which leads to increased latency and affects user experience.
Deep learning models are optimized through model optimization methods, including techniques such as model pruning, quantization, and distillation, to reduce storage requirements and computational complexity, and to adapt to the computing power and generalization capabilities of different devices.
It enables efficient deployment of deep learning models on resource-constrained devices, reduces storage requirements and computational complexity, and improves the computational efficiency and generalization ability of the models.
Smart Images

Figure CN121968137A_ABST
Abstract
Description
Model optimization methods, communication equipment and storage media Technical Field
[0001] This application relates to the field of communication technology, specifically to a model optimization method, communication equipment, and storage medium. Background Technology
[0002] Deep learning models typically contain a large number of parameters, such as large Transformer or CNN models, which may have hundreds of millions or even billions of parameters. These models require significant disk space for storage, making them difficult to deploy on resource-constrained devices. Furthermore, the inference process of deep learning models usually requires substantial computational resources, especially in scenarios with high real-time requirements (such as autonomous driving and real-time speech recognition). High computational complexity leads to increased latency, impacting user experience. In practical applications, models need to be deployed on various devices, including resource-constrained edge devices. Large models, due to their high storage requirements and computational complexity, are difficult to deploy directly on these devices. Summary of the Invention
[0003] In view of this, embodiments of this application provide a model optimization method, a communication device, and a storage medium, which achieve the technical effect of optimizing the model.
[0004] This application provides a model optimization method applied to a first network element node, including:
[0005] Receive a model operation request sent by a second network element node; wherein the model operation request includes model operation information;
[0006] Model optimization is performed based on the aforementioned model operation information.
[0007] This application provides a model optimization method applied to a second network element node, including:
[0008] A model operation request is sent to the first network element node so that the first network element node can perform model optimization based on the model operation information; wherein, the model operation request includes model operation information.
[0009] This application provides a model optimization device applied to a first network element node, comprising:
[0010] The receiving module is configured to receive a model operation request sent by the second network element node; wherein the model operation request includes model operation information.
[0011] The optimization module is configured to optimize the model based on the model operation information.
[0012] This application provides a model optimization device applied to a second network element node, comprising:
[0013] The sending module is configured to send a model operation request to a first network element node, so that the first network element node can perform model optimization based on the model operation information; wherein, the model operation request includes model operation information.
[0014] This application provides a communication device, including: a memory, and one or more processors;
[0015] The memory is configured to store one or more programs;
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the above embodiments.
[0017] This application provides a storage medium storing a computer program, which, when executed by a processor, implements the methods described in any of the above embodiments. Attached Figure Description
[0018] Figure 1 is a schematic diagram of a network management architecture provided by the prior art;
[0019] Figure 2 is a flowchart of a model optimization method provided in an embodiment of this application;
[0020] Figure 3 is a flowchart of another model optimization method provided in an embodiment of this application;
[0021] Figure 4 is an interactive schematic diagram of model optimization provided in an embodiment of this application;
[0022] Figure 5 is an interactive schematic diagram of another model optimization provided in an embodiment of this application;
[0023] Figure 6 is an interactive schematic diagram of another model optimization provided in an embodiment of this application;
[0024] Figure 7 is an interactive diagram of another model optimization provided in an embodiment of this application;
[0025] Figure 8 is an interactive diagram of another model optimization provided in an embodiment of this application;
[0026] Figure 9 is a structural block diagram of a model optimization device provided in an embodiment of this application;
[0027] Figure 10 is a structural block diagram of another model optimization device provided in an embodiment of this application;
[0028] Figure 11 is a schematic diagram of the structure of a communication device provided in an embodiment of this application. Detailed Implementation
[0029] The embodiments of this application will be described below with reference to the accompanying drawings. The examples given are for illustrative purposes only and are not intended to limit the scope of this application.
[0030] Figure 1 is a schematic diagram of a network management architecture provided by the prior art. As shown in Figure 1, the network management architecture may include a Business Support System (BSS), a Cross-Domain Management Function (CD-MnF), a Domain Management Function (Domain-MnF), and Network Elements (NEs). The CD-MnF manages one or more Domain Management Functions. The Domain Management Functions can manage one or more Network Elements.
[0031] A business support system (BSS) is oriented towards communication services and provides functions and management services such as billing, settlement, accounting, customer service, sales, network monitoring, communication service lifecycle management, and service intent translation. The BSS can be an operator's operating system or a vertical industry operating system (vertical OTSystem).
[0032] The Cross-Domain Management Function (NMF), also known as the Network Management Function (NMF), can be a network management entity such as a Network Management System (NMS), a Network Management Service Producer (MnS Producer), a Network Management Service Consumer (MnS Consumer), and a Network Function Management Service Consumer (NFMS_C). The NMF provides one or more of the following management functions or services: network lifecycle management, network deployment, network fault management, network performance management, network configuration management, network assurance, network optimization, and translation of network intents from Communication Service Providers (Intent-CSPs). The network referred to in the above management functions or services can include one or more network elements or subnetworks, or it can be a network slice. In other words, a network management function unit can be a Network Slice Management Function (NSMF), a Management Data Analytical Function (MDAF), a Self-Organization Network Function (SON), an Intent-Driven Management Service (Intent Driven MnS), or a Close Control Loop Management Service.
[0033] The domain management function unit, also known as the network subnet management function unit (NSMF) or network element management function unit, can be a wireless automation engine (MBB automationengine, MAE), a network element management system (EMS), a network function management service provider (NFMS_P), a network slice subnet management function (NSSMF), a domain management data analysis function (Domain MDAF), a self-organization network function (SON Function), a domain intent management function unit, or a close control loop management service, an MnS producer, an MnS consumer, or other network element management entities. Domain management function units can be classified in the following ways: By network type, they can be divided into: Radio Access Network (RAN) Domain Management Function (RAN domain MnF), Core Network Domain Management Function (CN domain MnF), and Transport Network Domain Management Function (TN domain MnF), etc. It should be noted that a domain management function unit can also be a domain network management system, managing one or more of the access network, core network, or transport network. By administrative region, they can be divided into: domain management function units for a specific region, such as the domain management function unit for city A, the domain management function unit for city B, etc.The domain management functional unit provides one or more of the following functions or management services: lifecycle management of subnetworks or network elements, deployment of subnetworks or network elements, fault management of subnetworks or network elements, performance management of subnetworks or network elements, assurance of subnetworks or network elements, optimization functions of subnetworks or network elements, and translation of subnetwork or network element intents (Intent from Network Operator, Intent-NOP), etc. Here, a subnetwork includes one or more network elements. A subnetwork can also include subnetworks, that is, one or more subnetworks forming a larger subnetwork. Here, a subnetwork can also be a network slice subnetwork.
[0034] A network element is an entity that provides network services, including core network elements, radio access network elements, or transport network elements. Specifically, core network elements may include, but are not limited to, Access and Mobility Management Function (AMF) entities, Session Management Function (SMF) entities, Policy Control Function (PCF) entities, Network Data Analysis Function (NWDAF) entities, Network Repository Function (NRF) entities, and gateways. Wireless access network elements may include, but are not limited to: various base stations (such as Generation NodeBs (gNBs), Evolved NodeBs (eNBs), Central Unit Control Panels (CUCPs), Central Units (CUs), Distributed Units (DUs), and Central Unit User Panels (CUUPs), etc. In this application, Network Functions (NFs) are also referred to as Network Elements (NEs). Network elements can provide one or more of the following management functions or services: network element lifecycle management, deployment, fault management, performance management, assurance, optimization functions, and translation of network element intents, etc.
[0035] This application proposes a model optimization method for use in networks, which addresses issues related to model storage, computational efficiency, deployment, energy efficiency, and generalization ability through various approaches. These methods can be used individually or in combination to achieve better compression results, making deep learning models more suitable for application in various real-world scenarios.
[0036] It should be noted that in this application, any consumer / producer can be a UE or any entity in Figure 1, namely a UE, gNB, NF (such as NWDAF), Operation Administration and Management (OAM) (such as MDAF), cross-domain OAM, Operations Support System (OSS) / BSS.
[0037] In one embodiment, Figure 2 is a flowchart of a model optimization method provided by an embodiment of this application. This embodiment is applied to the case of model optimization in a network. This embodiment can be executed by a first network element node. Exemplarily, the first network element node can be a Producer. As shown in Figure 2, this embodiment includes: S210-S220.
[0038] S210, Receive the model operation request sent by the second network element node; wherein, the model operation request contains model operation information.
[0039] In one example, a model operation request may include one of the following types: a model optimization request and a model acquisition request. In another example, the specific content of the model operation information depends on the type of the model operation request. In one example, if the model operation request is a model optimization request, the model operation information is model optimization information; that is, the model optimization request contains model optimization information. In another example, if the model operation request is a model acquisition request, the model operation information is the capability information of a second network element node or other third-party entities; that is, the model acquisition request contains capability information. For example, the second network element node can be a Consumer.
[0040] S220. Optimize the model based on model operation information.
[0041] In one example, a model optimization request can be sent directly through the second network element node to trigger the model optimization process to the first network element node. After the first network element node receives the model optimization request, it performs model optimization based on the model optimization information in the model optimization request.
[0042] In one example, the first network element node includes two different producers, such as a model training producer and a model optimization producer. The model optimization process can be triggered by the model training producer. That is, the second network element node can send a model acquisition request to the model training producer, so that the model training producer can determine whether model optimization is needed based on the capability information of the second network element node. If model optimization is needed, the model training producer sends a model optimization request to the model optimization producer, so that the model optimization producer can perform the model optimization process, so that the optimized model can achieve the effect required by the user, including performance, size and running requirements, and thus make the deep learning model suitable for application in various network scenarios.
[0043] In one embodiment, the model operation request includes: a model optimization request; the model operation information is model optimization information;
[0044] Accordingly, the first network element node includes either a model training producer or a model optimization producer; receiving model operation requests sent by the second network element node includes: receiving a model optimization request sent by the second network element node; wherein the model optimization request contains model optimization information. In one example, the model optimization process can be triggered by the second network element node: if the first network element node is a model training producer, the second network element node can send a model optimization request to the model training producer, and include model optimization information in the model optimization request; the model training producer performs model optimization based on the received model optimization request. In another example, the model optimization process can be triggered by the second network element node: if the first network element node is a model optimization producer, the second network element node can send a model optimization request to the model optimization producer, and include model optimization information in the model optimization request; the model optimization producer performs model optimization based on the received model optimization request.
[0045] In one embodiment, the model operation request includes: a model acquisition request; the model operation information includes capability information. In one example, the capability information can be the capability information of a second network element node, or it can be the capability information of other third-party entities. In one example, the capability information refers to the computational capability to perform model inference.
[0046] In one embodiment, the first network element node includes: a model training producer and a model optimization producer; receiving model operation requests sent by the second network element node includes: receiving a model acquisition request sent by the second network element node through the model training producer. In one example, a model acquisition request refers to a request message to obtain a model from the second network element node. The second network element node may send a model acquisition request to the model training producer, and include capability information in the model acquisition request. This capability information may be the capability information of the second network element node itself, or it may be the capability information of other third-party entities.
[0047] In one embodiment, the model optimization method applied to the first network element node further includes: a model training producer evaluating whether to perform model optimization and the model optimization category based on the capability information in the model acquisition request; if the model training producer evaluates that model optimization is necessary, it sends a model optimization request to the model optimization producer, or the model training producer performs model optimization based on the capability information evaluation result; wherein, the model optimization request includes: model optimization information. In one example, the first network element node may include a model training producer and a model optimization producer, and the model optimization process may be triggered by the model training producer: after the model training producer receives the model acquisition request sent by the second network element node, the model training producer can evaluate whether model optimization is needed and the model optimization category based on the capability information in the model acquisition request; if the model training producer determines that model optimization is needed, the model training producer can send a model optimization request to the model optimization producer, and the model optimization request includes model optimization information, so that the model optimization producer can perform the model optimization process based on the model optimization information. In one example, the capability information assessment result is used to characterize information to confirm whether model optimization is needed. It can also be understood as confirming model optimization information. For example, the capability information assessment result may include, but is not limited to, at least one of the following: confirming the model optimization category (i.e., confirming the model optimization category adopted by the model), confirming the model computational efficiency parameters (confirming the computational efficiency and computational capabilities that the model can support after optimization), and confirming the model generalization capability parameters (confirming the generalization capabilities that the model can support after optimization). The first network element node may include a model training producer and a model optimization producer, and the model optimization process can be triggered by the model training producer: After the model training producer receives a model acquisition request sent by the second network element node, the model training producer can assess whether model optimization is needed and the model optimization category based on the capability information in the model acquisition request. If the model training producer determines that model optimization is needed, it can directly execute the model optimization process based on the capability information assessment result.
[0048] In one embodiment, the first network element node includes: a model training producer and / or a model optimization producer; the model optimization method applied to the first network element node further includes: sending a model optimization report to a second network element node; wherein the model optimization report includes model optimization records. In one example, the model optimization records are used to record the model optimization process, the parameters of the model optimization, and the specific values of the model optimization information used, such as parameters included in the model optimization information, such as model optimization category, model optimization time, and optimized model performance. In one example, when the model optimization process is triggered by the model training producer, after the model optimization producer executes the model optimization process, the model optimization producer can send a model optimization report to the model training producer, and the model training producer forwards the model optimization report to the second network element node, so that the second network element node receives the optimized model and transmits or deploys the optimized model to the associated management entity. In one example, when the model optimization process is triggered by the second network element node and the model optimization producer executes the model optimization, after the model optimization producer completes the model optimization process, the model optimization producer can directly send a model optimization report to the second network element node. In one example, when the model optimization process is triggered by the second network element node and executed by the model training producer, the model training producer can directly send a model optimization report to the second network element node after completing the model optimization process.
[0049] In one embodiment, the capability information includes at least one of the following: computational efficiency parameters; infrastructure parameters. In one example, if the capability information is the capability information of a second network element node, the computational efficiency parameters are used to indicate the computational efficiency supported by the second network element node and the computational capabilities supported by the current hardware running model; for example, computational capabilities may include FLOPs, number of parameters, memory usage, inference time, inference speed, and supported model quantization accuracy, etc.; in one example, infrastructure parameters may include, but are not limited to, one of the following: CPU, GPU, and DPU models and capabilities.
[0050] In one embodiment, the model optimization information includes at least one of the following: model optimization indication; model optimization category; computational efficiency parameter; generalization ability parameter; model pruning information; model quantization information; and model distillation information. In one example, the model optimization indication is used to indicate model optimization, such as model compression.
[0051] In one embodiment, the model optimization category includes at least one of the following: model pruning; model quantization; and model distillation. In one example, the model optimization category indicates a specific model optimization method, and may include one or more of the following: model pruning, model quantization, and model distillation. In one example, model pruning and model quantization are commonly used model compression techniques, and model distillation is a model training technique for achieving knowledge transfer and can also be used for model optimization. In one example, model pruning, model quantization, and model distillation are all used to reduce the model's storage requirements and computational complexity while obtaining a model suitable for the consumer's computing power.
[0052] In one embodiment, computational efficiency parameters are used to indicate the computational efficiency and computational capabilities supported by the optimized model. These parameters include at least one of the following: floating-point operations per second (FLOPs); number of parameters; memory usage; inference time; inference speed; and supported model quantization precision. In one example, computational efficiency parameters indicate the computational efficiency that the optimized model needs to meet, as well as the hardware requirements, i.e., the computational capabilities required to run the model. These parameters can characterize the resource consumption and speed of the model during inference. For example, computational efficiency parameters may include, but are not limited to, at least one of the following: floating-point operations per second (FLOPs), number of parameters, memory usage, inference time, inference speed, and supported model quantization precision.
[0053] In one embodiment, the generalization ability parameter is used to indicate the generalization ability supported by the optimized model; the generalization ability parameter includes at least one of the following: accuracy; precision; recall; F1 score; test error; partial variance decomposition. In one example, generalization ability is used to measure the model's performance on untrained data, and for example, at least one of accuracy, precision, recall, F1 score, test error, and partial variance decomposition can be used to measure the model's generalization ability.
[0054] In one embodiment, the model pruning information includes at least one of the following: the original model; pruning configuration parameters; and pruning training data. In one example, the original model may include at least one of the following parameters for characterization: model identifier, model file, model address, model inference category (AIMLINferenceName, such as a specific MDA Type or Analitic ID), and model capability information (InferenceCapability, such as coverage inference). In one example, the pruning training data is used to fine-tune the pruned model to recover the performance lost due to pruning.
[0055] In one embodiment, the pruning configuration parameters include at least one of the following: pruning granularity; pruning ratio; pruning benchmark information; and pruning method. In one example, the pruning granularity may include at least one of the following: global or local, weighted pruning, channel pruning, and neuron pruning, etc.; in one example, the pruning ratio is used to characterize how many percentage of weights need to be retained; the pruning benchmark information refers to the pruning criteria, which can be selected using weight magnitude, gradient magnitude, or importance score; the pruning method may include one of the following: dynamic, static, and structured, etc.
[0056] In one embodiment, the model quantization information includes at least one of the following: the original model; quantization configuration parameters; and quantization calibration data. In one example, the quantization configuration parameters indicate the configuration and requirements of the specific quantization process; the quantization calibration data is used to determine a representative dataset for the quantization parameters (e.g., scaling factors and zero points), and the calibration data is typically a small set of input samples used to simulate the actual input to the model. It should be noted that the explanation of the original model included in the model quantization information can be found in the description of the original model included in the model pruning information above, and will not be repeated here.
[0057] In one embodiment, the quantization configuration parameters include at least one of the following: quantization precision; quantization method; and quantization mode. In one example, the quantization method may include, but is not limited to, at least one of the following: symmetric quantization, asymmetric quantization, uniform quantization, non-uniform quantization, static quantization, and dynamic quantization. In one example, the quantization precision may include 8-bit quantization (int8), 4-bit quantization (int4), 16-bit quantization (int16), and 32-bit quantization (int32), etc.; the quantization mode may include per-tensor or per-channel.
[0058] In one embodiment, the model distillation information includes at least one of the following: a first type of model; a second type of model; distillation configuration parameters; and distillation training data; wherein the performance of the first type of model is greater than that of the second type of model. In one example, the greater performance of the first type of model can also be understood as the stronger inference ability of the first type of model compared to the second type of model, or the greater inference range of the first type of model compared to the second type of model. In one example, the first type of model can be a pre-trained or pre-specialized large model with good performance, i.e., a pre-trained model or an augmented pre-trained model, which has a larger inference range; the second type of model can be a simpler and more computationally efficient model with a smaller inference range, focusing on one or more inference categories. The first type of model or the second type of model can be characterized by at least one of the following parameters: model identifier, model file, model inference category (AIMLINferenceName, such as a specific MDA Type or Analitic ID), and model capability information (InferenceCapability, such as coverageinference). In one example, distilled training data is used to train a second type of model and make it learn the output of a first type of model. Typically, the distilled training data is the same as the training data for the first type of model.
[0059] In one embodiment, the distillation configuration parameters include at least one of the following: a distillation loss function; and a distillation strategy. In one example, the distillation loss function typically includes knowledge distillation loss (such as KL divergence) and original task loss (such as cross-entropy loss); the distillation strategy may include one of the following: static distillation and dynamic distillation.
[0060] In one embodiment, Figure 3 is a flowchart of another model optimization method provided by an embodiment of this application. This embodiment is applied to the case of model optimization in a network. This embodiment can be executed by a second network element node. Exemplarily, the second network element node can be a Consumer. As shown in Figure 3, this embodiment includes:
[0061] S310. Send a model operation request to the first network element node so that the first network element node can perform model optimization based on the model operation information; wherein, the model operation request contains model operation information.
[0062] In one embodiment, the model operation request includes: a model optimization request; the model operation information is model optimization information;
[0063] Accordingly, the first network element node includes a model training producer or a model optimization producer; sending model operation requests to the first network element node includes:
[0064] Send a model optimization request to the first network element node; the model optimization request contains model optimization information.
[0065] In one embodiment, the model operation request includes: a model acquisition request; the model operation information includes capability information.
[0066] In one embodiment, the first network element node includes: a model training producer; and a model operation request sent to the first network element node, including:
[0067] Send a model retrieval request to the model training producer so that the model training producer can evaluate whether to optimize the model based on the capability information in the model retrieval request;
[0068] If model optimization is required, the model training producer sends a model optimization request to the model optimization producer, or the model training producer optimizes the model based on the capability information evaluation results; wherein, the model optimization request contains model optimization information.
[0069] In one embodiment, the first network element node includes: a model training producer and / or a model optimization producer; the model optimization method applied to the second network element node further includes:
[0070] Receive the model optimization report from the first network element node; the model optimization report includes: model optimization records.
[0071] In one embodiment, the capability information includes at least one of the following: computational efficiency parameters; infrastructure parameters.
[0072] In one embodiment, the model optimization information includes at least one of the following: model optimization indication; model optimization category; computational efficiency parameter; generalization ability parameter; model pruning information; model quantization information; and model distillation information.
[0073] In one embodiment, the model optimization category includes at least one of the following: model pruning; model quantization; model distillation.
[0074] In one embodiment, the computational efficiency parameter is used to indicate the computational efficiency and computational capability supported by the optimized model; the computational efficiency parameter includes at least one of the following: floating-point operations per second; number of parameters; memory usage; inference time; inference speed; supported model quantization precision.
[0075] In one embodiment, the generalization ability parameter is used to indicate the generalization ability supported by the optimized model; the generalization ability parameter includes at least one of the following: accuracy; precision; recall; F1 score; test error; partial variance decomposition.
[0076] In one embodiment, the model pruning information includes at least one of the following: the original model; pruning configuration parameters; and pruning training data.
[0077] In one embodiment, the pruning configuration parameters include at least one of the following: pruning granularity; pruning ratio; pruning reference information; and pruning method.
[0078] In one embodiment, the model quantization information includes at least one of the following: the original model; quantization configuration parameters; and quantization calibration data.
[0079] In one embodiment, the quantization configuration parameters include at least one of the following: quantization accuracy; quantization method; quantization mode.
[0080] In one embodiment, the model distillation information includes at least one of the following: a first type of model; a second type of model; distillation configuration parameters; and distillation training data; wherein the model performance of the first type of model is greater than that of the second type of model.
[0081] In one embodiment, the distillation configuration parameters include at least one of the following: a distillation loss function; a distillation strategy.
[0082] It should be noted that the explanations of parameters such as model operation request, model operation information, model optimization request, and model optimization information in the model optimization method applied to the second network element node can be found in the descriptions of the corresponding parameters in the model optimization method applied to the first network element node, and will not be repeated here.
[0083] It should be noted that the Producer, consumer, and Entity in the following embodiments can all correspond to the entities in Figure 1, namely cross-domain OAM, single-domain OAM, gNB, NWDAF, or other network elements.
[0084] Model pruning, model quantization, and model distillation are several commonly used model optimization or compression techniques, all aiming to reduce model storage requirements and computational complexity while improving inference speed. The model optimization provided in this application includes two scenarios: first, consumer-triggered; second, producer-triggered.
[0085] Example 1
[0086] Figure 4 is an interactive schematic diagram of model optimization provided in an embodiment of this application. In this embodiment, taking the first network element node as the Producer, the second network element node as the Consumer, the first type of model as the teacher model, the second type of model as the student model, the model operation request as the model optimization request, and the model operation information as the model optimization information, the model optimization process of the Consumer request is explained. As shown in Figure 4, the model optimization process of this embodiment includes S410-S430.
[0087] S410. The Consumer sends a model optimization request to the Producer; the model optimization request contains model optimization information.
[0088] In one example, model optimization information includes one of the following: model optimization indication, model optimization category, computational efficiency parameters, generality parameters, model pruning information, model quantization information, and model distillation information.
[0089] In one example, the model optimization instruction indicates that model optimization should be performed, including model compression, etc.
[0090] Model optimization category: indicates the specific model optimization method, which may include one or more of the following: model pruning, model quantization, and model distillation;
[0091] Computational efficiency parameters: These indicate the computational efficiency that the optimized model needs to meet, the hardware requirements, and the computational capabilities that the model needs to support. They mainly characterize the resource consumption and speed of the model during inference, such as FLOPs, number of parameters, memory usage, inference time, inference speed, and supported model quantization accuracy.
[0092] Generalization ability parameter: indicates the generalization ability that the optimized model should meet. Generalization ability measures the model's performance on unseen data, such as accuracy, precision, recall, F1 score, test error, partial variance decomposition, etc.
[0093] In one example, model pruning information (applicable when the model optimization category is model pruning) includes:
[0094] Original model: such as original model file, original model identifier;
[0095] Pruning configuration parameters include: pruning granularity (global or local, weight pruning, channel pruning, neuron pruning, etc.), pruning ratio (e.g., how many percentage of weights to retain), pruning baseline information (also known as pruning standard information, such as weight size, gradient size, or importance score to select the pruning target), and pruning method (dynamic, static, structured).
[0096] Pruned training data: Used to fine-tune the pruned model to recover performance that may have been lost due to pruning.
[0097] In one example, model quantization information (applicable when the model optimization category is model quantization) includes:
[0098] Original model: such as original model file, original model identifier;
[0099] Quantization configuration parameters: These indicate the specific configuration and requirements of the quantization process, including quantization precision (e.g., int8 or int4), quantization method (symmetric quantization, asymmetric quantization, static quantization, dynamic quantization), and quantization mode (e.g., per-tensor or per-channel).
[0100] Quantization calibration data: A representative dataset used to determine quantization parameters (such as scaling factors and zeros). Calibration data is typically a small set of input samples used to simulate the actual inputs of the model.
[0101] In one example, model distillation information (applicable when the model optimization category is model distillation) includes:
[0102] The first type of model, also known as the teacher model, is a large, pre-trained model with good performance.
[0103] The second type of model, also known as the student model, is a model with a simpler structure and higher computational efficiency.
[0104] Distillation training data: Used to train the student model to learn the output of the teacher model, and is usually the same as the training data of the teacher model.
[0105] Distillation configuration parameters: Distillation loss functions typically include knowledge distillation loss (such as KL divergence) and original task loss (such as cross-entropy loss); distillation strategies, such as static distillation and dynamic distillation.
[0106] S420, the Producer performs model optimization based on the model optimization information in the model optimization request.
[0107] S430, the Producer sends a model optimization report to the Consumer; the model optimization report includes model optimization records.
[0108] Example 2
[0109] Figure 5 is an interactive schematic diagram of another model optimization provided in this embodiment. In this embodiment, taking the first network element node as the Producer, the second network element node as the Consumer, the model operation request as the model acquisition request, and the model operation information as capability information as an example, the model optimization process triggered by the Producer request is explained. As shown in Figure 5, the model optimization process in this embodiment includes S510-S530. As shown in Figure 5, the model optimization process in this embodiment includes the following steps:
[0110] S510. The Consumer sends a model retrieval request to the Producer, which includes the Consumer's capability information.
[0111] S520: The Producer assesses whether model optimization is needed and the type of model optimization based on the capability information carried in the model acquisition request.
[0112] In one example, the Consumer's capability information, which indicates the Consumer's computational capabilities for performing model inference, may include the following:
[0113] Computational efficiency parameters: These indicate the computational efficiency supported by the consumer, the computational capabilities supported by the current hardware model running on the consumer, such as FLOPs, number of parameters, memory usage, inference time, inference speed, and supported model quantization accuracy.
[0114] Infrastructure parameters: such as CPU / GPU / DPU models and capabilities.
[0115] In one example, if the Producer needs to optimize the model based on capability information assessment, the Producer sends a model optimization request to other network element nodes so that the other network element nodes can perform model optimization.
[0116] S530, the Producer sends a model optimization report, which includes model optimization records.
[0117] Example 3
[0118] Figure 6 is a schematic diagram of another model optimization interaction provided in an embodiment of this application. In this embodiment, taking the first network element node as the ML model training producer and the second network element node as the ML model training consumer, the first type of model as the teacher model and the second type of model as the student model, the model operation request as the model optimization request, and the model operation information as the model optimization information, the model optimization process of the consumer request is explained. This embodiment can be understood as a further detailed explanation of the model optimization scheme shown in Figure 4, and the model optimization is performed by the ML Model Training Producer. As shown in Figure 6, the model optimization process of this embodiment includes S610-S680.
[0119] S610, Perform a query for the ML Model Training Producer.
[0120] The ML Model Training Consumer can send a query request for the ML Model Training Producer to the Registration & Discovery Producer. The Registration & Discovery Producer responds with the ML Model Training Producer's Domain Name (DN).
[0121] S620: Send a model training request to the ML Model Training Producer.
[0122] The ML Model Training Consumer sends a model training request to the ML Model Training Producer, which includes the model's AIMLInferenceName. The ML Model Training Producer responds with an ML Model.
[0123] S630, ML Model Training Consumer queries the capability information of managed objects.
[0124] Capability information, used to indicate the Consumer's computational capabilities for performing model inference, may include the following:
[0125] Computational efficiency parameters: These indicate the computational efficiency supported by the consumer, the computational capabilities supported by the current hardware model running on the consumer, such as FLOPs, number of parameters, memory usage, inference time, inference speed, and supported model quantization accuracy.
[0126] Infrastructure parameters: such as CPU / GPU / DPU models and capabilities.
[0127] S640. Evaluate whether model optimization is necessary.
[0128] The ML Model Training Consumer evaluates whether model optimization is needed based on the received model and capability information.
[0129] S650: Send a model optimization request to the ML Model Training Consumer.
[0130] The ML Model Training Consumer sends a model optimization request to the ML Model Training Producer. The model optimization request includes at least one of the following parameters: model optimization instruction, model optimization category, computational efficiency parameter, generality parameter, model pruning information, model quantization information, and model distillation information.
[0131] Model optimization information may include the following:
[0132] Model optimization instructions: Instructions to perform model optimization, including model compression, etc.
[0133] Model optimization category: indicates the specific model optimization method, which may include one or more of the following: model pruning, model quantization, model distillation;
[0134] Computational efficiency parameters: These indicate the computational efficiency that the optimized model needs to meet, the hardware requirements, and the computational capabilities that the model needs to support. They mainly characterize the resource consumption and speed of the model during inference, such as FLOPs, number of parameters, memory usage, inference time, inference speed, and supported model quantization accuracy.
[0135] Generalization ability parameter: indicates the generalization ability that the optimized model should meet. Generalization ability measures the model's performance on unseen data, such as accuracy, precision, recall, F1 score, test error, partial variance decomposition, etc.
[0136] Model pruning information (applicable when the model optimization category is model pruning) may include:
[0137] Original model: such as original model file, original model identifier;
[0138] Pruning configuration parameters: pruning granularity (global or local, weight pruning, channel pruning, neuron pruning, etc.), pruning ratio (e.g., what percentage of weights to retain), pruning criteria (selecting pruning targets based on weight size, gradient size, or importance score), pruning method (dynamic, static, structured), etc.
[0139] Pruned training data: Used to fine-tune the pruned model to recover performance that may have been lost due to pruning.
[0140] Model quantization information (applicable when the model optimization category is model quantization) may include:
[0141] Original model: such as original model file, original model identifier;
[0142] Quantization configuration parameters: These indicate the specific configuration and requirements of the quantization process, including quantization precision (e.g., int8 or int4), quantization method (symmetric quantization, asymmetric quantization, static quantization, dynamic quantization), and quantization mode (e.g., per-tensor or per-channel).
[0143] Quantization calibration data: A representative dataset used to determine quantization parameters (such as scaling factors and zeros). Calibration data is typically a small set of input samples used to simulate the actual inputs of the model.
[0144] Model distillation information (applicable when the model optimization category is model distillation) may include:
[0145] Teacher Model: A pre-trained, high-performance large model;
[0146] Student Model: A model with a simpler structure and higher computational efficiency;
[0147] Distillation training data: Used to train the student model to learn the output of the teacher model, and is usually the same as the training data of the teacher model;
[0148] Distillation configuration parameters: Distillation loss functions typically include knowledge distillation loss (such as KL divergence) and original task loss (such as cross-entropy loss); distillation strategies, such as static distillation and dynamic distillation.
[0149] S660, perform model optimization based on model optimization information.
[0150] The ML Model Training Producer optimizes the model based on the received model optimization requests.
[0151] S670, Send model optimization report.
[0152] The ML Model Training Producer sends a model optimization report to the ML Model Training Consumer. This report includes model optimization records, which document information about the optimization process, such as the model optimization category used and the model optimization parameters.
[0153] S680: Transmit and deploy the optimized model.
[0154] The ML Model Training Producer transfers or deploys optimized models to Managed Entities.
[0155] Example 4
[0156] Figure 7 is a schematic diagram of another model optimization interaction provided in an embodiment of this application. In this embodiment, taking the first network element node as including an ML Model Training Producer and an ML Model Optimization Producer; the second network element node as an ML Model Training Consumer; the first type of model as a teacher model; the second type of model as a student model; the model operation request as a model optimization request; and the model operation information as model optimization information, the model optimization process of the Consumer request is explained. This embodiment can be understood as a further detailed explanation of the model optimization scheme shown in Figure 4, and the model optimization is performed by the ML Model Optimization Producer. As shown in Figure 7, the model optimization process in this embodiment includes S710-S790.
[0157] S710, perform a query for ML Model Training Producer.
[0158] The ML Model Training Consumer can send a query request to the Registration & Discovery Producer for the MLTraining Producer. The Registration & Discovery Producer responds with the DN of the ML ModelTraining Producer.
[0159] S720: Send a model training request to the ML Model Training Producer.
[0160] The ML Model Training Consumer sends a model training request to the ML Model Training Producer, which includes the model's AIMLInferenceName. The ML Model Training Producer responds with an ML Model.
[0161] S730, ML Model Training Consumer queries the capability information of managed objects.
[0162] Capability information, used to indicate the Consumer's computational capabilities for performing model inference, may include the following:
[0163] Computational efficiency parameters: These indicate the computational efficiency supported by the consumer, the computational capabilities supported by the current hardware model running on the consumer, such as FLOPs, number of parameters, memory usage, inference time, inference speed, and supported model quantization accuracy.
[0164] Infrastructure parameters: such as CPU / GPU / DPU models and capabilities.
[0165] S740. Assess whether model optimization is necessary.
[0166] The ML Training Consumer assesses whether model optimization is needed based on the received model and capability information.
[0167] S750, query ML Model Optimization Producer.
[0168] Consumer queries Registration&Discovery Producer for ML Model OptimizationProducer.
[0169] S760: Send a model optimization request to the ML Model Optimization Producer.
[0170] The ML Model Training Consumer sends a model optimization request to the ML Model Optimization Producer. The model optimization request includes at least one of the following parameters: model optimization instruction, model optimization category, computational efficiency parameter, generality parameter, model pruning information, model quantization information, and model distillation information.
[0171] Model optimization information may include the following:
[0172] Model optimization instructions: Instructions to perform model optimization, including model compression, etc.
[0173] Model optimization category: indicates the specific model optimization method, which may include one or more of the following: model pruning, model quantization, model distillation;
[0174] Computational efficiency parameters: These indicate the computational efficiency that the optimized model needs to meet, the hardware requirements, and the computational capabilities that the model needs to support. They mainly characterize the resource consumption and speed of the model during inference, such as FLOPs, number of parameters, memory usage, inference time, inference speed, and supported model quantization accuracy.
[0175] Generalization ability parameter: indicates the generalization ability that the optimized model should meet. Generalization ability measures the model's performance on unseen data, such as accuracy, precision, recall, F1 score, test error, partial variance decomposition, etc.
[0176] Model pruning information (applicable when the model optimization category is model pruning) may include:
[0177] Original model: such as original model file, original model identifier;
[0178] Pruning configuration parameters: pruning granularity (global or local, weight pruning, channel pruning, neuron pruning, etc.), pruning ratio (e.g., what percentage of weights to retain), pruning criteria (selecting pruning targets based on weight size, gradient size, or importance score), pruning method (dynamic, static, structured), etc.
[0179] Pruned training data: Used to fine-tune the pruned model to recover performance that may have been lost due to pruning.
[0180] Model quantization information (applicable when the model optimization category is model quantization) may include:
[0181] Original model: such as original model file, original model identifier;
[0182] Quantization configuration parameters: These indicate the specific configuration and requirements of the quantization process, including quantization precision (e.g., int8 or int4), quantization method (symmetric quantization, asymmetric quantization, static quantization, dynamic quantization), and quantization mode (e.g., per-tensor or per-channel).
[0183] Quantization calibration data: A representative dataset used to determine quantization parameters (such as scaling factors and zeros). Calibration data is typically a small set of input samples used to simulate the actual inputs of the model.
[0184] Model distillation information (applicable when the model optimization category is model distillation) may include:
[0185] Teacher Model: A large, well-performing, pre-trained or pre-specialized model, i.e., a pre-trained model or an augmented pre-trained model, which has a larger inference range.
[0186] Student Model: A model with a simpler structure and higher computational efficiency, with a smaller reasoning scope, focusing on one or more reasoning categories;
[0187] Distillation training data: Used to train the student model to learn the output of the teacher model, and is usually the same as the training data of the teacher model;
[0188] Distillation configuration parameters: Distillation loss functions typically include knowledge distillation loss (such as KL divergence) and original task loss (such as cross-entropy loss); distillation strategies, such as static distillation and dynamic distillation.
[0189] S770, perform model optimization based on model optimization information.
[0190] The ML Model Optimization Producer optimizes the model based on the received model optimization requests.
[0191] S780, Send model optimization report.
[0192] The ML Model Optimization Producer sends a model optimization report to the ML Model Training Consumer. This report includes model optimization records, which document information about the optimization process, such as the model optimization category used and the model optimization parameters.
[0193] S790: Transmit and deploy the optimized model.
[0194] The ML Model Training Producer transfers or deploys optimized models to Managed Entities.
[0195] It should be noted that the differences between this fourth embodiment and the third embodiment described above include:
[0196] Firstly, the S750 has been added, allowing the Consumer to query the ML ModelOptimization Producer from the Registration & Discovery Producer.
[0197] Secondly, the difference between S760-S780 and S650-S670 in Example 3 is that the main body performing model optimization changes from ML Model Training Producer to ML Model Optimization Producer.
[0198] Example 5
[0199] Figure 8 is a schematic diagram of another model optimization interaction provided in an embodiment of this application. In this embodiment, taking the first network element node as including an ML Model Training Producer and an ML Model Optimization Producer; the second network element node as an ML Model Training Consumer; the first type of model as a teacher model; the second type of model as a student model; the model operation request as a model acquisition request; and the model operation information as capability information as an example, the model optimization process requested by the ML ModelTraining Producer is explained. This embodiment can be understood as a further detailed explanation of the model optimization scheme shown in Figure 5, and the model optimization is performed by the ML Model Optimization Producer. As shown in Figure 8, the model optimization process in this embodiment includes S810-S880.
[0200] S810, perform a query for ML Model Training Producer.
[0201] The ML Model Training Consumer sends a query request to the Registration & Discovery Producer for the ML Model Training Producer. The Registration & Discovery Producer responds with the DN of the ML ModelTraining Producer.
[0202] S820: Send a model retrieval request to the ML Model Training Producer.
[0203] The ML Model Training Consumer sends a model retrieval request to the ML Model Training Producer. This request includes the model's AIMLInferenceName and the Consumer's capability information. Here, "consumer" can be the Consumer that issued the model retrieval request (i.e., the ML Model Training Consumer) or the capability information of other third-party entities.
[0204] S830. Based on the Consumer's capability information, assess whether model optimization is needed and the type of model optimization.
[0205] The ML Model Training Producer assesses whether model optimization is needed and the type of optimization based on the Consumer's capability information.
[0206] S840, Send model optimization request.
[0207] The ML Model Training Producer sends a model optimization request, which includes model optimization information, to the ML Model Optimization Producer.
[0208] S850, Execution model optimization.
[0209] The ML Model Optimization Producer performs model optimization.
[0210] S860, Send model optimization report.
[0211] The ML Model Optimization Producer sends a model optimization report to the ML Model Training Producer.
[0212] The S870 ML Model Training Producer sends model optimization reports, including model optimization records, to the ML Model Training Consumer.
[0213] In one embodiment, FIG9 is a structural block diagram of a model optimization device provided in this application. This embodiment is applied to a first network element node. As shown in FIG9, the model optimization device in this embodiment includes a receiving module 910 and an optimization module 920.
[0214] The receiving module 910 is configured to receive a model operation request sent by the second network element node; wherein the model operation request contains model operation information.
[0215] Optimization module 920 is configured to perform model optimization based on model operation information.
[0216] In one embodiment, the model operation request includes: a model optimization request; the model operation information is model optimization information;
[0217] Accordingly, the first network element node includes a model training producer or a model optimization producer; it receives model operation requests sent by the second network element node, including:
[0218] Receive the model optimization request sent by the second network element node; the model optimization request contains model optimization information.
[0219] In one embodiment, the model operation request includes: a model acquisition request; the model operation information includes capability information.
[0220] In one embodiment, the first network element node includes: a model training producer and a model optimization producer; receiving model operation requests sent by the second network element node includes:
[0221] The model training producer receives model acquisition requests sent by the second network element node.
[0222] In one embodiment, the model optimization device applied to the first network element node further includes:
[0223] The evaluation module is configured so that the model training producer evaluates whether to perform model optimization and the type of model optimization based on the capability information in the model acquisition request.
[0224] The sending module is configured to send a model optimization request to the model optimization producer if the model training producer evaluates and performs model optimization, or the model training producer performs model optimization based on the capability information evaluation results; wherein, the model optimization request includes: model optimization information.
[0225] In one embodiment, the first network element node includes: a model training producer and a model optimization producer; the model optimization device applied to the first network element node further includes:
[0226] The sending module is also configured to send a model optimization report to the second network element node; wherein, the model optimization report includes: model optimization records.
[0227] In one embodiment, the capability information includes at least one of the following: computational efficiency parameters; infrastructure parameters.
[0228] In one embodiment, the model optimization information includes at least one of the following: model optimization indication; model optimization category; computational efficiency parameter; generalization ability parameter; model pruning information; model quantization information; and model distillation information.
[0229] In one embodiment, the model optimization category includes at least one of the following: model pruning; model quantization; model distillation.
[0230] In one embodiment, the computational efficiency parameter is used to indicate the computational efficiency and computational capability supported by the optimized model; the computational efficiency parameter includes at least one of the following: floating-point operations per second; number of parameters; memory usage; inference time; inference speed; supported model quantization precision.
[0231] In one embodiment, the generalization ability parameter is used to indicate the generalization ability supported by the optimized model; the generalization ability parameter includes at least one of the following: accuracy; precision; recall; F1 score; test error; partial variance decomposition.
[0232] In one embodiment, the model pruning information includes at least one of the following: the original model; pruning configuration parameters; and pruning training data.
[0233] In one embodiment, the pruning configuration parameters include at least one of the following: pruning granularity; pruning ratio; pruning reference information; and pruning method.
[0234] In one embodiment, the model quantization information includes at least one of the following: the original model; quantization configuration parameters; and quantization calibration data.
[0235] In one embodiment, the quantization configuration parameters include at least one of the following: quantization accuracy; quantization method; quantization mode.
[0236] In one embodiment, the model distillation information includes at least one of the following: a first type of model; a second type of model; distillation configuration parameters; and distillation training data; wherein the model performance of the first type of model is greater than that of the second type of model.
[0237] In one embodiment, the distillation configuration parameters include at least one of the following: a distillation loss function; a distillation strategy.
[0238] The model optimization device provided in this embodiment is configured to implement the model optimization method applied to the first network element node in the embodiment shown in Figure 2. The implementation principle and technical effect of the model optimization device provided in this embodiment are similar, and will not be described again here.
[0239] In one embodiment, FIG10 is a structural block diagram of another model optimization device provided in this application embodiment. This embodiment is applied to a second network element node. As shown in FIG10, the model optimization device in this embodiment includes: a sending module 1010.
[0240] The sending module 1010 is configured to send a model operation request to the first network element node so that the first network element node can perform model optimization based on the model operation information; wherein, the model operation request contains model operation information.
[0241] In one embodiment, the model operation request includes: a model optimization request; the model operation information is model optimization information;
[0242] Accordingly, the first network element node includes a model training producer or a model optimization producer; sending model operation requests to the first network element node includes:
[0243] Send a model optimization request to the first network element node; the model optimization request contains model optimization information.
[0244] In one embodiment, the model operation request includes: a model acquisition request; the model operation information includes capability information.
[0245] In one embodiment, the first network element node includes: a model training producer; and a model operation request sent to the first network element node, including:
[0246] Send a model retrieval request to the model training producer so that the model training producer can evaluate whether to optimize the model based on the capability information in the model retrieval request;
[0247] If model optimization is to be performed, the model training producer sends a model optimization request to the model optimization producer, or the model training producer performs model optimization based on the capability information evaluation results; wherein, the model optimization request includes model optimization information.
[0248] In one embodiment, the first network element node includes: a model training producer and a model optimization producer; the model optimization device applied to the second network element node further includes:
[0249] The receiving module is configured to receive the model optimization report fed back by the first network element node; wherein, the model optimization report includes: model optimization record.
[0250] In one embodiment, the capability information includes at least one of the following: computational efficiency parameters; infrastructure parameters.
[0251] In one embodiment, the model optimization information includes at least one of the following: model optimization indication; model optimization category; computational efficiency parameter; generalization ability parameter; model pruning information; model quantization information; and model distillation information.
[0252] In one embodiment, the model optimization category includes at least one of the following: model pruning; model quantization; model distillation.
[0253] In one embodiment, the computational efficiency parameter is used to indicate the computational efficiency and computational capability supported by the optimized model; the computational efficiency parameter includes at least one of the following: floating-point operations per second; number of parameters; memory usage; inference time; inference speed; supported model quantization precision.
[0254] In one embodiment, the generalization ability parameter is used to indicate the generalization ability supported by the optimized model; the generalization ability parameter includes at least one of the following: accuracy; precision; recall; F1 score; test error; partial variance decomposition.
[0255] In one embodiment, the model pruning information includes at least one of the following: the original model; pruning configuration parameters; and pruning training data.
[0256] In one embodiment, the pruning configuration parameters include at least one of the following: pruning granularity; pruning ratio; pruning reference information; and pruning method.
[0257] In one embodiment, the model quantization information includes at least one of the following: the original model; quantization configuration parameters; and quantization calibration data.
[0258] In one embodiment, the quantization configuration parameters include at least one of the following: quantization accuracy; quantization method; quantization mode.
[0259] In one embodiment, the model distillation information includes at least one of the following: a first type of model; a second type of model; distillation configuration parameters; and distillation training data; wherein the model performance of the first type of model is greater than that of the second type of model.
[0260] In one embodiment, the distillation configuration parameters include at least one of the following: a distillation loss function; a distillation strategy.
[0261] The model optimization device provided in this embodiment is configured to implement the model optimization method applied to the second network element node in the embodiment shown in Figure 3. The implementation principle and technical effect of the model optimization device provided in this embodiment are similar, and will not be described again here.
[0262] In one embodiment, FIG11 is a schematic diagram of the structure of a communication device provided in an embodiment of this application. As shown in FIG11, the device provided in this application includes: a processor 1110, a memory 1120, and a communication module 1130. The number of processors 1110 in the device can be one or more; FIG11 shows one processor 1110 as an example. The number of memories 1120 in the device can be one or more; FIG11 shows one memory 1120 as an example. The processor 1110, memory 1120, and communication module 1130 of the device can be connected via a bus or other means; FIG11 shows a connection via a bus as an example. In this embodiment, the device can be a first network element node or a second network element node.
[0263] The memory 1120, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the device in any embodiment of this application (e.g., the receiving module 910 and optimization module 920 in the model optimization apparatus). The memory 1120 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application program required for at least one function; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory 1120 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 1120 may further include memory remotely located relative to the processor 1110, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0264] When the communication device is the first network element node, the device provided above can be configured to execute the model optimization method applied to the first network element node provided in any of the above embodiments, and has the corresponding functions and effects.
[0265] When the communication device is a second network element node, the device provided above can be configured to execute the model optimization method applied to the second network element node provided in any of the above embodiments, and has the corresponding functions and effects.
[0266] This application embodiment also provides a storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to perform a model optimization method applied to a first network element node. The method includes: receiving a model operation request sent by a second network element node; wherein the model operation request includes model operation information; and performing model optimization based on the model operation information.
[0267] This application embodiment also provides a storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to execute a model optimization method applied to a second network element node. The method includes: sending a model operation request to a first network element node so that the first network element node performs model optimization based on model operation information; wherein the model operation request contains model operation information.
[0268] Those skilled in the art will understand that the term user equipment covers any suitable type of wireless user equipment, such as mobile phones, portable data processing devices, portable web browsers, or vehicle-mounted mobile stations.
[0269] Generally, the various embodiments of this application can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, although this application is not limited thereto.
[0270] Embodiments of this application can be implemented by executing computer program instructions through the data processor of a mobile device, for example, in a processor entity, or through hardware, or through a combination of software and hardware. The computer program instructions can be assembly instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages.
[0271] Any block diagram of logical flow in the accompanying drawings of this application may represent program steps, or may represent interconnected logic circuits, modules, and functions, or may represent a combination of program steps and logic circuits, modules, and functions. The computer program may be stored on memory. Memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as, but not limited to, read-only memory (ROM), random access memory (RAM), optical storage devices and systems (Digital Video Disc (DVD) or Compact Disk (CD)), etc. Computer-readable media may include non-transitory storage media. The data processor may be of any type suitable to the local technical environment, such as, but not limited to, general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and processors based on multi-core processor architectures.
[0272] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A model optimization method, characterized in that, Applied to a first network element node, the method includes: receiving a model operation request sent by a second network element node; wherein the model operation request includes model operation information; and performing model optimization based on the model operation information.
2. The method according to claim 1, characterized in that, The model operation request includes a model optimization request; the model operation information is model optimization information; correspondingly, the first network element node includes a model training producer or a model optimization producer; receiving the model operation request sent by the second network element node includes receiving the model optimization request sent by the second network element node; wherein, the model optimization request includes model optimization information.
3. The method according to claim 1, characterized in that, The model operation request includes: a model acquisition request; the model operation information includes capability information.
4. The method according to claim 3, characterized in that, The first network element node includes: a model training producer and a model optimization producer; receiving the model operation request sent by the second network element node includes: receiving the model acquisition request sent by the second network element node through the model training producer.
5. The method according to claim 4, characterized in that, The method further includes: the model training producer assessing whether to perform model optimization and the model optimization category based on the capability information in the model acquisition request; if the model training producer assesses that model optimization is necessary, it sends a model optimization request to the model optimization producer, or the model training producer performs model optimization based on the capability information assessment result; wherein, the model optimization request includes: model optimization information.
6. The method according to claim 2 or 4, characterized in that, The first network element node includes: a model training producer and / or a model optimization producer; the method further includes: sending a model optimization report to the second network element node; wherein the model optimization report includes: model optimization records.
7. The method according to claim 3, characterized in that, The capability information includes at least one of the following: computational efficiency parameters; infrastructure parameters.
8. The method according to claim 2 or 5, characterized in that, The model optimization information includes at least one of the following: model optimization indication; model optimization category; computational efficiency parameters; generalization ability parameters; model pruning information; model quantization information; and model distillation information.
9. The method according to claim 8, characterized in that, The model optimization category includes at least one of the following: model pruning; model quantization; model distillation.
10. The method according to claim 8, characterized in that, The computational efficiency parameters are used to indicate the computational efficiency and computational capabilities supported by the optimized model; the computational efficiency parameters include at least one of the following: floating-point operations per second; number of parameters; memory usage; inference time; inference speed; and supported model quantization precision.
11. The method according to claim 8, characterized in that, The generalization ability parameter is used to indicate the generalization ability supported by the optimized model; the generalization ability parameter includes at least one of the following: accuracy; precision; recall; F1 score; test error; partial variance decomposition.
12. The method according to claim 8, characterized in that, The model pruning information includes at least one of the following: the original model; pruning configuration parameters; and pruning training data.
13. The method according to claim 12, characterized in that, The pruning configuration parameters include at least one of the following: pruning granularity; pruning ratio; pruning reference information; pruning method.
14. The method according to claim 8, characterized in that, The model quantization information includes at least one of the following: the original model; quantization configuration parameters; and quantization calibration data.
15. The method according to claim 14, characterized in that, The quantization configuration parameters include at least one of the following: quantization accuracy; quantization method; quantization mode.
16. The method according to claim 8, characterized in that, The model distillation information includes at least one of the following: a first type of model; a second type of model; distillation configuration parameters; distillation training data; wherein the model performance of the first type of model is greater than that of the second type of model.
17. The method according to claim 16, characterized in that, The distillation configuration parameters include at least one of the following: distillation loss function; distillation strategy.
18. A model optimization method, characterized in that, The method is applied to a second network element node, including: sending a model operation request to a first network element node so that the first network element node performs model optimization based on the model operation information; wherein the model operation request includes model operation information.
19. The method according to claim 18, characterized in that, The model operation request includes: a model optimization request; the model operation information is model optimization information; correspondingly, the first network element node includes a model training producer or a model optimization producer; sending the model operation request to the first network element node includes: sending the model optimization request to the first network element node; wherein, the model optimization request contains model optimization information.
20. The method according to claim 18, characterized in that, The model operation request includes: a model acquisition request; the model operation information includes capability information.
21. The method according to claim 20, characterized in that, The first network element node includes: a model training producer; the model operation request sent to the first network element node includes: sending a model acquisition request to the model training producer, so that the model training producer evaluates whether to perform model optimization based on the capability information in the model acquisition request; if model optimization is performed, the model training producer sends a model optimization request to the model optimization producer, or the model training producer performs model optimization based on the capability information evaluation result; wherein, the model optimization request contains model optimization information.
22. The method according to claim 18, characterized in that, The first network element node includes: a model training producer and / or a model optimization producer; the method further includes: receiving a model optimization report fed back by the first network element node; wherein the model optimization report includes: model optimization records.
23. A communication device, characterized in that, include: Memory, and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors perform the method as described in any one of claims 1-17 or 18-22.
24. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-17 or 18-22.