Model generation method and device, communication equipment and readable storage medium

By dynamically generating small models that are adapted to services on the AI RAN side, the problem of limited computing resources is solved, and efficient inference accuracy and business quality assurance are achieved.

CN120302304APending Publication Date: 2025-07-11CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510384862.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

AI RAN计算资源受限,难以支持复杂的大模型推理,计算精度下降,影响业务质量。

Method used

When a preset trigger condition is monitored, the server sends a model update request to the central cloud, and the central cloud responds and distiles knowledge based on the large model to generate a small model, deploys it on the access network, and dynamically adapts to the computing requirements.

Benefits of technology

It reduces the computing requirements of AI RAN, ensures inference accuracy, and improves business quality and computing resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302304A_ABST
    Figure CN120302304A_ABST
Patent Text Reader

Abstract

The invention relates to a model generation method and device, communication equipment and a computer readable storage medium. The method comprises the following steps: when a preset triggering condition is monitored, sending a model updating request to a central cloud; the central cloud responds to the model updating request and determines a small model according to a pre-deployed large model; and receiving the small model returned by the central cloud, and deploying the small model in an access network. By adopting the method, the calculation requirement of the AI RAN can be reduced, and the problem that calculation resources are limited is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wireless communication technologies, and particularly to a model generation method, apparatus, communication device, and computer-readable storage medium. Background Art

[0002] Artificial Intelligence Radio Access Network (AI RAN) is a network architecture that combines artificial intelligence with a radio access network, supporting intelligent radio management and optimizing computing resource scheduling to provide efficient and intelligent services between terminal devices and the network.

[0003] In traditional technologies, large-scale AI models such as Generative Pre-trained Transformer (GPT) models can be directly deployed and run on the AI RAN side. This method relies on the computing resources provided by an edge computing platform. Combining technologies such as model pruning and quantization can further reduce the computing overhead. In addition, by using a cache and computing sharing mechanism to multiplex multi-user tasks on the AI RAN side, duplicate computations can also be reduced. However, due to limited computing resources in the AI RAN, even with technologies such as pruning and quantization, it is still difficult to support complex large model inferences, and due to computing power limitations, the inference accuracy may decrease, affecting service quality.

[0004] Therefore, there is a problem of limited computing resources in current AI RAN technologies. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a model generation method, apparatus, communication device, computer-readable storage medium, and computer program product that can avoid limited computing resources.

[0006] In a first aspect, this application provides a model generation method applied to a server, including:

[0007] When a preset trigger condition is monitored, send a model update request to the central cloud; the central cloud responds to the model update request and determines a small model according to a pre-deployed large model;

[0008] Receive the small model returned by the central cloud and deploy the small model in the access network.

[0009] In a second aspect, this application provides a model generation method applied to the central cloud, including:

[0010] In response to the model update request from the server, determine a small model according to the pre-deployed large model;

[0011] Return the small model to the server; the server deploys the small model in the access network.

[0012] In a third aspect, the present application provides a model generation method, which is applied to the access network and includes:

[0013] Obtain a small model; the small model is determined by the cloud according to the pre-deployed large model and is obtained after the server deploys the determined small model;

[0014] Infer task data according to the small model to obtain an inference result of the task data.

[0015] In a fourth aspect, the present application further provides a model generation device, which is applied to the server and includes:

[0016] A sending module, configured to send a model update request to the central cloud when a preset trigger condition is detected; the central cloud responds to the model update request and determines a small model according to the pre-deployed large model;

[0017] A receiving module, configured to receive the small model returned by the central cloud and deploy the small model in the access network.

[0018] In a fifth aspect, the present application further provides a communication device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the first aspect, the second aspect, and the third aspect are implemented.

[0019] In a sixth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the first aspect, the second aspect, and the third aspect are implemented.

[0020] In a seventh aspect, the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the method described in any one of the first aspect, the second aspect, and the third aspect are implemented.

[0021] The above model generation method, apparatus, communication device, computer-readable storage medium, and computer program product enable the server to send a model update request to the central cloud when a preset trigger condition is detected. The central cloud responds to the model update request, determines a small model based on a pre-deployed large model, and the server receives the small model returned by the central cloud and deploys the small model in the access network. When the preset trigger condition is met, the central cloud can be requested to perform knowledge distillation on the pre-deployed large model to dynamically generate a small model adapted to AI RAN for AI RAN to perform inference, reducing the computational requirements of AI RAN and solving the problem of limited computing resources. Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0023] Figure 1 It is an application environment diagram of the model generation method in an embodiment;

[0024] Figure 2 It is a flowchart of the model generation method in an embodiment;

[0025] Figure 3 It is a schematic diagram of the wireless intelligent management and orchestration function in an embodiment;

[0026] Figure 4 It is an interaction flowchart of the intelligent scheduling method for an automatic production robot based on AI RAN in an embodiment;

[0027] Figure 5 It is an interaction flowchart of the model generation method in an embodiment;

[0028] Figure 6 It is a structural block diagram of the model generation apparatus in an embodiment;

[0029] Figure 7 It is an internal structure diagram of the communication device in an embodiment. Detailed Embodiments

[0030] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0031] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0032] The model generation method provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the access network 102 communicates with the server 104 through the network, and the server 104 communicates with the central cloud 106 through the network. Among them, the access network 102 can be, but is not limited to, an AI RAN, which is connected to the terminal and provides inference services, task scheduling, etc. for the terminal. The central cloud (Central Cloud) 106 can be, but is not limited to, a cloud server with large-scale computing power and storage capacity. The server 104 can be a core network or an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a control and scheduling device (RAN AI Layer) of the AI RAN. This control and scheduling device can be set separately or integrated on other devices for controlling and scheduling the AI RAN.

[0033] In an exemplary embodiment, as Figure 2 shown, a model generation method is provided. In this embodiment, an example is given in which this method is applied to the Figure 1 server 104 in, including the following steps:

[0034] Step 202, when a preset trigger condition is detected, send a model update request to the central cloud; the central cloud responds to the model update request and determines a small model according to the pre-deployed large model.

[0035] Among them, the preset trigger condition can be that the existing small model of the access network does not meet the service requirements. For example, the calculation accuracy is lower than the preset threshold or the resource utilization rate is lower than the preset threshold. The preset trigger condition can also be that the performance of the large model in the central cloud is upgraded, and a small model with better performance can be provided for the access network.

[0036] Among them, the model update request can be a message requesting to update the small model of the access network. The large model can be an AI model with relatively more parameters, higher computing resource requirements and stronger performance. The small model can be an AI model with relatively fewer parameters, lower computing resource requirements and moderate performance.

[0037] In a specific implementation, the server can monitor the access network or the central cloud. When it detects that the preset trigger condition is met, it sends a model update request to the central cloud. When the central cloud receives the model update request, it can perform knowledge distillation on the large model pre-deployed on itself to obtain a small model.

[0038] For example, when the AI RAN uses an existing small model to perform task scheduling for a terminal, the RAN AI Layer can monitor the current network environment, computing resource status, and service requirements, collect terminal data, and analyze whether the existing small model on the AI RAN side needs to be optimized based on the network environment, computing resource status, service requirements, and terminal data. If the existing small model on the AI RAN side cannot meet the service requirements (such as insufficient computing accuracy or low resource utilization), it means that optimization is needed. At this time, the RAN AI Layer can send a model update request to the central cloud and carry the current service information on the AI RAN side (such as task type, computing resource status, data characteristics, etc.). After receiving the model update request, the central cloud performs knowledge distillation on the large model pre-deployed on itself according to the current service information on the AI RAN side to obtain a small model adapted to the current service information on the AI RAN side.

[0039] Step 204: Receive the small model returned by the central cloud and deploy the small model in the access network.

[0040] In a specific implementation, after distilling a small model from the large model, the central cloud can send the small model to the server, and the server deploys the received small model in the access network.

[0041] For example, the central cloud can send the small model generated by knowledge distillation to the RAN AI Layer, and the RAN AI Layer deploys the small model sent by the central cloud in the AI RAN. It can be understood that the RAN AI Layer can be connected to one or more AI RANs. When multiple different AI RANs all need to update the small model, the server can deploy multiple small models on the corresponding AI RANs respectively.

[0042] In the above model generation method, when the server detects the preset trigger condition, it sends a model update request to the central cloud. The central cloud responds to the model update request, determines the small model according to the pre-deployed large model, the server receives the small model returned by the central cloud and deploys the small model in the access network; it can request the central cloud to perform knowledge distillation on the pre-deployed large model when the preset trigger condition is met, dynamically generate a small model adapted to the AI RAN for the AI RAN to perform inference, reduce the computing requirements of the AI RAN, and solve the problem of limited computing resources.

[0043] In practical applications, directly applying the small model distilled from the large model to the access network may result in the inference accuracy of the small model not meeting expectations. Therefore, the small model can be continuously optimized until its inference accuracy meets the expectations. In an exemplary embodiment, after the above step 204, it may further specifically include: receiving the first inference result of the small model sent by the access network; forwarding the first inference result to the central cloud; the central cloud compares the first inference result with the second inference result of the large model to obtain the inference accuracy of the small model. If the inference accuracy is lower than the preset threshold, an optimization signal is returned; when the optimization signal is received, the small model of the access network is updated.

[0044] Among them, the first inference result can be the inference result of the small model. The second inference result can be the inference result of the large model. The inference accuracy can be the accuracy of the inference result. The preset threshold can be a preset inference accuracy threshold. The optimization signal can be a signal indicating to optimize the small model.

[0045] In specific implementation, after deploying the small model in the access network, the access network can use the current small model to perform inference on the task data to obtain the first inference result, and send the first inference result and the task data to the server. The server forwards the first inference result and the task data to the central cloud. The central cloud uses the large model to perform inference on the task data of the access network to obtain the second inference result, and compares the received first inference result with the second inference result generated by itself to obtain the inference accuracy of the small model. If the inference accuracy is lower than the preset threshold, it means that the inference accuracy of the current small model in the access network fails to meet expectations. At this time, the central cloud can generate an optimization signal and send the optimization signal to the server. The server updates the current small model in the access network when it receives the optimization signal; otherwise, if the inference accuracy is not lower than the preset threshold, it means that the inference accuracy of the current small model in the access network meets expectations and there is no need to update. At this time, the current small model in the access network can run independently without the need for the central cloud to cooperate and confirm.

[0046] Among them, the task data can be data related to the inference task to be executed. For example, in the intelligent production scenario, the task data can be part status, assembly accuracy, welding quality, etc. At this time, the corresponding inference results can be material classification, part matching, welding error, etc. In addition, it can be understood that let the first inference result be , and the second inference result be , then the calculation formula for the inference accuracy can be .

[0047] It should be noted that the central cloud can also forward the second inference result of the large model to the access network in real time through the server, so that the access network can perform task scheduling according to the second inference result. Compared with performing task scheduling using the first inference result, using the second inference result for task scheduling in the access network has higher accuracy.

[0048] It is understandable that the process of determining the inference accuracy based on the inference results of the large and small models, and then determining whether to continue optimizing the small model according to the inference accuracy can also be executed on the server or the access network. This application does not limit this.

[0049] For example, the access network uses the small model to infer the task data, obtains the first inference result, sends the first inference result and the task data to the server, the server forwards the task data to the central cloud, the central cloud uses the large model to infer the task data, and returns the obtained second inference result to the server. The server determines the inference accuracy of the small model based on the first inference result sent by the access network and the second inference result sent by the central cloud, and determines whether to continue optimizing the small model according to the inference accuracy.

[0050] Another example is that the access network uses the small model to infer the task data, obtains the first inference result, sends the task data to the server, the server forwards the task data to the central cloud, the central cloud uses the large model to infer the task data, obtains the second inference result, and forwards the second inference result to the access network through the server. The access network determines the inference accuracy of the small model based on the first inference result obtained by itself and the second inference result forwarded by the server, and determines whether to continue optimizing the small model according to the inference accuracy.

[0051] In this embodiment, by receiving the first inference result of the small model sent by the access network, forwarding the first inference result to the central cloud, the central cloud compares the first inference result with the second inference result of the large model to obtain the inference accuracy of the small model. In the case where the inference accuracy is lower than the preset threshold, an optimization signal is returned. When the optimization signal is received, the small model of the access network is updated, so that the small model can be continuously optimized when the inference accuracy of the small model fails to reach the expectation, and the accuracy of the small model inference is improved.

[0052] In practical applications, when the server receives the optimization signal, the small model can be updated by adjusting the model parameters. Therefore, in an exemplary embodiment, the above step of updating the small model of the access network may specifically include: adjusting the parameters of the small model to obtain new parameters of the small model; sending the new parameters to the access network; and the access network updating the small model according to the new parameters.

[0053] In specific implementation, when the optimization signal is received, the server can adjust the parameters of the current small model to obtain new parameters of the current small model, send the new parameters to the access network, and the access network updates the current small model according to the new parameters to obtain a more adaptable new small model.

[0054] In this embodiment, by adjusting the parameters of the small model, new parameters of the small model are obtained, and the new parameters are sent to the access network. The access network updates the small model according to the new parameters, which can quickly update the small model and improve the efficiency of model generation.

[0055] It can be understood that the update method of the small model is not unique. When the server receives an optimization signal, it can also instruct the central cloud to re-distill the small model. In an exemplary embodiment, the above steps of updating the small model of the access network may specifically further include: sending a model update request to the central cloud; the central cloud responds to the model update request and determines a new small model according to the large model; deploying the new small model returned by the central cloud in the access network.

[0056] In specific implementation, when receiving an optimization signal, the server can send a model update request to the central cloud again. When the central cloud receives the model update request, it can perform knowledge distillation on the large model again to obtain a new small model. The central cloud sends the new small model to the server, and the server deploys the new small model in the access network.

[0057] In this embodiment, by sending a model update request to the central cloud, the central cloud responds to the model update request, determines a new small model according to the large model, and deploys the new small model returned by the central cloud in the access network, which can improve the accuracy of the small model and further improve the accuracy of task scheduling in the access network.

[0058] It can be understood that when the central cloud compares the first inference result of the access network with its own second inference result and the obtained inference accuracy is not lower than the preset threshold, it means that the inference accuracy of the current small model can meet the requirements, and the access network can run the current small model independently without relying on the collaborative confirmation of the central cloud. At this time, the inference data generated during the independent operation of the access network can be synchronized to the central cloud, and the central cloud uses the inference data of the access network to optimize the performance of its own large model. Therefore, in an exemplary embodiment, the above model generation method may specifically further include: receiving the inference data uploaded by the access network; when the inference accuracy of the small model in the access network is not lower than the preset threshold, using the small model for inference and counting the inference data during the inference process of the small model; forwarding the inference data to the central cloud; the central cloud updates the large model according to the inference data.

[0059] Among them, the inference data may be the data generated during the process of the access network using the small model for inference. For example, in an intelligent production scenario, the task data may be part status, assembly accuracy, welding quality, etc., and the corresponding inference results are material classification, part matching, welding error, etc., then the inference data may be the task success rate, execution time, abnormal conditions, etc. generated during the inference process.

[0060] In a specific implementation, when the inference accuracy of the small model is not lower than a preset threshold, the access network can use the small model to perform inference on task data to obtain an inference result. The access network can count the inference data generated during the inference process of the small model and forward the inference data to the central cloud through the server. The central cloud updates its large model according to the received inference data.

[0061] In practical applications, the access network can synchronize the task data, inference data, inference results, etc. of the current small model to the central cloud. The central cloud can use these data as new training data to continuously train the large model and continuously optimize the performance of the large model. The central cloud can also distill a new small model from the optimized large model and feedback it to the server for updating the small model of the access network, thereby forming a closed-loop optimization mechanism for the small model.

[0062] In this embodiment, by receiving the inference data uploaded by the access network, when the inference accuracy of the small model is not lower than the preset threshold, the access network uses the small model for inference, counts the inference data during the inference process of the small model, and forwards the inference data to the central cloud. The central cloud updates the large model according to the inference data, which can continuously optimize the large model and the small model and improve the accuracy of task scheduling.

[0063] In an exemplary embodiment, the model update request carries the service information of the access network, and the central cloud performs knowledge distillation on the large model according to the service information to obtain a small model.

[0064] Among them, the service information can be data related to the user services carried by the access network, such as task type, computing resource status, data characteristics, etc.

[0065] In a specific implementation, the server can send a model update request to the central cloud. The model update request carries the service information of the access network. After receiving the model update request, the central cloud can perform knowledge distillation on the large model according to the service information in the model update request to obtain a small model.

[0066] For example, when it is detected that the small model on the AI RAN side needs to be optimized, the RAN AI Layer sends a model update request to the central cloud. The model update request carries the current service information (such as task type, computing resource status, data characteristics, etc.). After receiving the model update request, the central cloud uses the knowledge distillation technology to dynamically generate an adapted small model from the pre-deployed large model according to the service information on the AI RAN side.

[0067] In this embodiment, by carrying the service information of the access network in the model update request, the central cloud performs knowledge distillation on the large model according to the service information to obtain a small model, and can dynamically generate a lightweight small model adapted to the access network based on the large model, reducing the computing requirements of the access network while ensuring the inference accuracy and avoiding limited computing resources.

[0068] In an exemplary embodiment, a model generation method is provided. This embodiment takes the central cloud 106 in Figure 1 as an example for illustration, including the following steps:

[0069] Step 302, in response to the model update request from the server, determine the small model according to the pre-deployed large model;

[0070] Step 304, return the small model to the server; the server deploys the small model in the access network.

[0071] In a specific implementation, the server can monitor the access network or the central cloud. When it detects that the preset trigger condition is met, it sends a model update request to the central cloud. When the central cloud receives the model update request, it can perform knowledge distillation on the pre-deployed large model by itself to obtain a small model, send the small model to the server, and the server deploys the received small model in the access network.

[0072] In the above model generation method, by responding to the model update request from the server, determining the small model according to the pre-deployed large model, and returning the small model to the server, and the server deploys the small model in the access network; it can request the central cloud to perform knowledge distillation on the pre-deployed large model when the preset trigger condition is met, dynamically generate a small model adapted to the AI RAN for the AI RAN to perform inference, reduce the computing requirements of the AI RAN, and solve the problem of limited computing resources.

[0073] In an exemplary embodiment, the above step 302 may specifically include: obtaining the service information of the access network carried in the model update request; performing knowledge distillation on the large model according to the service information to obtain a small model.

[0074] In a specific implementation, the server can send a model update request to the central cloud, and the model update request carries the service information of the access network. After receiving the model update request, the central cloud can perform knowledge distillation on the large model according to the service information in the model update request to obtain a small model.

[0075] In this embodiment, by obtaining the service information of the access network carried in the model update request and performing knowledge distillation on the large model according to the service information, a small model can be obtained. A lightweight small model adapted to the access network can be dynamically generated based on the large model, reducing the computing requirements of the access network while ensuring the inference accuracy and avoiding limited computing resources.

[0076] In an exemplary embodiment, after the above step 304, it may further specifically include: receiving the first inference result of the access network forwarded by the server; comparing the first inference result with the second inference result of the large model to obtain the inference accuracy of the small model; and returning an optimization signal to the server when the inference accuracy is lower than the preset threshold.

[0077] In specific implementation, after the small model is deployed in the access network, the access network can use the current small model to perform inference on the task data to obtain a first inference result, and send the first inference result and the task data to the server. The server forwards the first inference result and the task data to the central cloud. The central cloud uses the large model to perform inference on the task data of the access network to obtain a second inference result, and compares the received first inference result with the second inference result generated by itself to obtain the inference accuracy of the small model. If the inference accuracy is lower than the preset threshold, it means that the inference accuracy of the current small model in the access network fails to meet the expectation. At this time, the central cloud can generate an optimization signal and send the optimization signal to the server. The server updates the current small model in the access network when receiving the optimization signal; otherwise, if the inference accuracy is not lower than the preset threshold, it means that the inference accuracy of the current small model in the access network meets the expectation and there is no need to update. At this time, the current small model in the access network can operate independently without the need for the central cloud to cooperate and confirm.

[0078] In this embodiment, by receiving the first inference result of the access network forwarded by the server, comparing the first inference result with the second inference result of the large model to obtain the inference accuracy of the small model, and returning an optimization signal to the server when the inference accuracy is lower than the preset threshold, the small model can be continuously optimized when the inference accuracy of the small model fails to meet the expectation, improving the accuracy of small model inference.

[0079] In an exemplary embodiment, a model generation method is provided. This embodiment takes this method applied to Figure 1 the access network 102 in as an example for illustration, including the following steps:

[0080] Step 402, obtain a small model; the small model is determined by the cloud according to the pre-deployed large model, and the server deploys the determined small model;

[0081] Step 404, perform inference on the task data according to the small model to obtain the inference result of the task data.

[0082] In a specific implementation, the server can monitor the access network or the central cloud. When it detects that the preset trigger condition is met, it sends a model update request to the central cloud. When the central cloud receives the model update request, it can perform knowledge distillation on the large model pre-deployed on itself to obtain a small model, and send the small model to the server. The server deploys the received small model in the access network, and the access network can perform inference on the task data according to the small model to obtain an inference result.

[0083] The above model generation method obtains a small model, which is determined by the cloud based on the pre-deployed large model, and is obtained by the server deploying the determined small model. The task data is inferred according to the small model to obtain the inference result of the task data. When the preset trigger condition is met, the central cloud can be requested to perform knowledge distillation on the pre-deployed large model, dynamically generate a small model adapted to AI RAN for AI RAN to perform inference, reduce the computing requirements of AI RAN, and solve the problem of limited computing resources.

[0084] To facilitate those skilled in the art to deeply understand the embodiments of the present application, the following will be described with a specific example.

[0085] Aiming at the problem that AI RAN has limited computing resources and it is difficult to directly run large models, the present application proposes a cooperative task scheduling method for the central cloud and AI RAN based on model distillation (Model Distillation). Its core goal is to achieve cloud-edge collaborative optimization of intelligent task scheduling, dynamically distill small models and adaptively deploy them on the AI RAN side to meet the requirements of different business scenarios. Specifically, the present application adopts a dynamic model distillation mechanism, enabling AI RAN to request the cloud large model for knowledge distillation when needed, and dynamically generating an adapted small model to reduce computing requirements while ensuring high-precision inference. Compared with the traditional offline distillation mechanism, this method can adjust the parameters of the small model in real time on the AI RAN side, enabling it to adapt to different business scenarios and improving the generalization ability of the model.

[0086] In addition, the present application realizes intelligent task scheduling through the RAN AI Layer. Figure 3A schematic diagram of the wireless intelligent management and orchestration function is provided. Among them, the wireless intelligent management and orchestration function is the RAN AI Layer. 3GPP RAN can be regarded as AI RAN, and cloud-native RAN can be regarded as the central cloud. The RAN AI Layer includes four core modules: model management function, data management service, service management orchestration, and communication and computing resource scheduling. It can intelligently manage the push of models, task execution, and resource scheduling on the AI RAN side to ensure the efficient coordination of model optimization and inference tasks in AI RAN. Among them, the model management function is responsible for the update of large models and small models, as well as parameter adjustment, etc. The data management service is responsible for functions such as data collection, reporting, and processing. The service management orchestration is responsible for managing the user services carried by the access network. The communication and computing resource scheduling is responsible for implementing the scheduling of communication resources and computing resources to jointly complete functions such as data transmission, reporting, and processing.

[0087] This application also proposes a cloud-edge closed-loop optimization system. In the initial stage of the operation of the AI RAN small model, the inference result will be compared and verified with the large model in the cloud. After ensuring that the inference result of the small model reaches an accuracy of 95% or so, the AI RAN small model can perform independent inference. This closed-loop optimization mechanism can ensure that the small model of AI RAN always maintains high accuracy, while reducing the dependence on the central cloud, improving the inference efficiency and network reliability.

[0088] The above-mentioned collaborative task scheduling method for the central cloud and AI RAN based on model distillation combines the powerful computing power of the cloud and fully utilizes the distributed inference ability of AI RAN, thus reducing the bandwidth occupancy while ensuring the inference accuracy, improving the independent operation ability of AI RAN, and ensuring the efficient and intelligent development of the wireless communication network in the AI era.

[0089] The above-mentioned collaborative task scheduling method for the central cloud and AI RAN based on model distillation mainly involves four core components: the central cloud, AIRAN, RAN AI Layer, and user equipment (UE).

[0090] Among them, the central cloud has large-scale computing and storage capabilities, and undertakes the training, optimization, and inference tasks of large models (such as multi-modal Transformer, generative AI, etc.); interacts with the RAN AI Layer through the API interface, responds to the model distillation request from the RAN AI Layer, and dynamically generates small models; has data analysis and closed-loop optimization functions, and can continuously adjust model parameters according to the feedback of AI RAN to improve the AI inference accuracy.

[0091] AI RAN has a certain computing power (integrating heterogeneous computing units such as CPUs, GPUs, and NPUs), but the accuracy of large models with large operating parameters is limited; it relies on the distilled small models provided by the central cloud and performs local inference, model caching, and edge data processing during business operations; it adjusts the operation strategy of the small model through wireless-side intelligent perception (such as base station channel state information, terminal behavior patterns, etc.).

[0092] The RAN AI Layer, as the core control and scheduling layer of AI RAN, is responsible for model management, data management, service orchestration, and computing resource scheduling; by monitoring the results of the operation of AI RAN in real time, it decides when to request the distilled small model from the cloud and evaluates the inference results of the small model in real time; it has the ability of cross-domain optimization and can coordinate the load balancing of AI tasks among multiple computing nodes of AI RAN.

[0093] The UE, which is a terminal device (such as a smart phone, VR / AR device, robot, etc.), accesses AI RAN wirelessly and utilizes the AI inference services provided by it; the business requirements of the UE (such as video stream analysis, speech recognition, etc.) trigger the inference calculation of the small model on the AI RAN side.

[0094] In one embodiment, a collaborative task scheduling method for the central cloud and AI RAN based on model distillation is proposed, aiming to solve the problems of difficult deployment of large models on the AI RAN side, high inference latency, poor model adaptability, and lack of cloud-edge closed-loop optimization. By constructing a dynamic collaboration mechanism between the cloud-based large model and the AI RAN small model, this method can efficiently deploy lightweight models on the AI RAN side, and while ensuring high inference accuracy, reduce the computing burden and improve the business response speed. This method mainly includes the following steps:

[0095] Step 501, RAN AI Layer task initialization and business requirement analysis. The RAN AI Layer in the control plane of AI RAN monitors the current network environment, computing resource situation, and business requirements, and collects user-side data through the data management module to analyze whether the current AI task needs to be optimized. If the existing small model on the AI RAN side cannot meet the business requirements (such as insufficient computing accuracy or low resource utilization rate), a dynamic model distillation request is triggered.

[0096] Step 502, Dynamic Model Distillation Request and Large Model Inference. When the AI RAN side needs to optimize the small model, the RANAI Layer sends a request to the large model in the central cloud (referred to as the cloud for short), carrying the current service feature information (such as task type, computing resource status, data features, etc.). After receiving the request, the cloud large model uses its powerful computing power for inference and returns high-precision inference results. At the same time, the cloud dynamically generates an adapted small model according to the computing power and service requirements of the AI RAN side using model distillation technology.

[0097] Step 503, Deployment and Initial Inference of the Small Model. The central cloud sends the distilled small model to the RAN AI Layer and automatically deploys it through the model management function module. At this stage, the small model on the AI RAN side does not operate independently but is in a stage that requires central cloud coordination and confirmation, that is, the inference results of the small model will be sent to the central cloud for comparison simultaneously to verify whether its inference accuracy meets the requirements (such as 95%).

[0098] Step 504, Cloud-Edge Closed-Loop Optimization Mechanism. If the inference accuracy of the small model on the AI RAN side fails to meet the expectation, the central cloud will feedback an optimization signal and adjust the parameters of the small model or re-distill a new small model through the model management function module until the small model on the AI RAN side reaches the set accuracy threshold. When the inference results of the small model are stable and the error is extremely small, the small model on the AI RAN side can operate independently, reducing dependence on the cloud, thereby improving inference efficiency and reducing bandwidth consumption.

[0099] Step 505, Continuous Model Optimization and Resource Scheduling. Using the general computing resource scheduling module, dynamically adjust the deployment strategy of the small model according to the computing power, service requirements, and network load conditions on the AI RAN side. For example, more efficient small models can be adopted during peak periods to reduce computing pressure, while during low loads, the model complexity can be appropriately increased to improve inference accuracy. In addition, the data on the AI RAN side can be continuously transmitted back to the central cloud for training more efficient distillation models, forming a long-term evolving optimization mechanism.

[0100] In this embodiment, through the dynamic model distillation and cloud-edge closed-loop optimization mechanism, the problem that the large model cannot be directly deployed on the AI RAN side is solved, and the computing efficiency is improved. In the traditional method, due to limited computing resources on the AI RAN, it is difficult to directly run the large model. In this application, through dynamic model distillation, the large model is only trained and run in the cloud, and an adapted small model is generated according to the requirements, enabling the AI RAN side to efficiently run AI inference tasks in a low-computing-resource environment.

[0101] Secondly, the problem of insufficient inference accuracy of small models is solved, and the intelligent level of AI RAN is improved. The cloud-edge collaborative closed-loop optimization mechanism proposed in this application can ensure that the inference accuracy of small models on the AI RAN side gradually approaches that of large models, reaching an accuracy of 95%, for example, to ensure service quality.

[0102] Thirdly, the problem of the disconnection between AI RAN task scheduling and model management is solved, and intelligent scheduling is realized. Through the four core modules of the RAN AI Layer (model management, data management, service orchestration, and communication computing resource scheduling), the integrated optimization of cloud-edge tasks is realized, enabling AI RAN to dynamically adjust model and resource configurations according to service requirements and improving network utilization.

[0103] In addition, the problem of high cloud-edge collaborative inference latency is solved, and bandwidth consumption is reduced. Through the dynamic deployment and independent inference mechanism of small models, this application reduces the dependence on the cloud on the AI RAN side, reduces bandwidth consumption, and improves inference efficiency. At the same time, in the case of unstable networks, AI RAN can still operate independently, improving system reliability.

[0104] Finally, the problem of long model optimization cycles is solved, and an adaptive optimization ability is constructed. Through the continuous data feedback and dynamic distillation mechanism, this application forms a cloud-edge closed-loop optimization system, enabling the small models on the AI RAN side to continuously evolve, adapt to new service requirements, and maintain optimal performance in the long term.

[0105] In one embodiment, an intelligent scheduling method for an automatic production robot based on AI RAN is proposed, which is applicable to the automatic production line of an intelligent manufacturing factory, such as Figure 4 shown. In this scenario, the robot needs to perform complex tasks such as visual recognition, fine operation, and quality inspection, which require high computing power. However, the traditional local computing method results in high equipment costs, high power consumption, and difficulty in real-time model optimization. This method adopts a collaborative computing architecture of the central cloud, RAN AI Layer, and AI RAN, and realizes efficient and low-latency intelligent control of the robot by dynamically distilling large models to the edge. The method mainly includes the following steps:

[0106] Step 601, Production environment data collection. After the production task is started, the automated production robot enters the working area and begins to execute the data collection task. The high-definition camera, infrared sensor, and force feedback sensor on the robot capture the environmental information on the production line in real time, including part status, assembly accuracy, welding quality, etc. Subsequently, the robot uploads this raw data to the AI RAN, and the AI RAN uses the small model deployed locally to perform the first-round inference on the data, initially judging key parameters such as material classification, part matching, and welding error. The results of the AI RAN inference will be synchronized to the data management module of the RAN AI Layer, and the RAN AI Layer further compares the historical data and adjusts the model parameters in combination with the real-time production status to ensure the accuracy and consistency of the production data.

[0107] Step 602, Model distillation and inference optimization. When the AI RAN completes the preliminary inference, its results will be uploaded to the RAN AI Layer for accuracy verification. If it is found that the inference error of the small model of the AI RAN is relatively large (below the set accuracy standard of 99.9999%), the RAN AI Layer will request the central cloud to re-distill the small model. After receiving the request, the central cloud uses the large model for optimization and adjustment, generates a new small model, and distributes it to the RAN AI Layer again, which is then distributed to the AI RAN for replacement. At the same time, the RAN AI Layer recalculates the priority of the production task based on the optimized model and adjusts the scheduling strategy of the robot task to ensure the optimal production rhythm. Through this dynamic optimization process, the small model of the AI RAN can gradually improve the inference accuracy, reduce the dependence on the central cloud, and improve the task execution efficiency.

[0108] Step 603, Robot task execution and dynamic adjustment. When the inference accuracy of the small model of the AI RAN reaches the expected standard, the robot can execute the task independently. The AI RAN assigns specific operations to the robot according to the task instructions issued by the RAN AI Layer, such as grasping components, precision assembly, welding, quality inspection, etc. While executing the task, the robot uses the small model provided by the AI RAN for real-time visual analysis to ensure the accuracy of the operation. For example, during the welding process, if the robot detects a deviation in the weld seam, the AI RAN can immediately adjust the welding parameters and optimize the welding path to ensure that the welding quality meets the standard. If the robot encounters an abnormal situation (such as component misalignment, material shortage) during the task execution, the AI RAN will immediately trigger the real-time adjustment mechanism, modify the execution parameters or re-plan the task path to reduce production errors and improve production continuity.

[0109] Step 604, Data Feedback and Continuous Optimization. After the production task is completed, AI RAN will automatically summarize the data generated during the robot's execution, including task success rate, execution time, abnormal situations, etc., and upload it to the data management module of the RAN AI Layer. The RAN AI Layer analyzes based on this data, optimizes the production scheduling strategy, and adjusts the inference process of AI RAN to make it more suitable for the current production environment. At the same time, key data will be synchronized to the central cloud. The central cloud uses the new data for continuous training, continuously optimizes the performance of the large model, and updates the small model of AI RAN through model distillation technology, forming a complete closed-loop optimization mechanism. Through this process, the robot can not only operate efficiently in the current production task but also continuously improve its intelligence level in future tasks, making the production process more accurate, stable, and efficient.

[0110] The above intelligent scheduling method for the automatic production robot based on AI RAN realizes the intelligent distillation of the large model and the dynamic deployment of the small model through the collaborative optimization of the central cloud and AI RAN. This method introduces a large model supervision mechanism at the initial stage of inference. After ensuring that the accuracy of the small model's inference results reaches the 95% standard, it is allowed to run independently. This progressive optimization strategy effectively reduces the inference error, improves the stability of the robot's task execution, and reduces the bandwidth consumption of cloud inference.

[0111] Secondly, as the intelligent management plane of AI RAN, the RAN AI Layer integrates the model management, data management, business management orchestration, and computing resource scheduling modules, making the model deployment no longer a static one-time operation but able to be dynamically adapted according to the computing power of AI RAN, the current task requirements, and the network load situation. This method can automatically adjust the deployment strategy of the small model according to the production environment and task requirements, thereby optimizing the computing resource allocation, avoiding computing power waste, and improving the overall throughput capacity of the system.

[0112] Thirdly, a dynamic distillation optimization mechanism is introduced, enabling AI RAN not only to execute tasks but also to continuously optimize based on real-time feedback data. This method transmits the AI RAN data to the central cloud in real time, and the central cloud continuously trains the large model based on the feedback data and regularly distills and pushes the optimized small model to AI RAN, realizing a self-learning and adaptive closed-loop optimization to ensure that the small model always has the optimal inference ability.

[0113] Finally, by distilling on demand and dynamically updating the small model, AI RAN can select the most suitable model version according to the computing resource situation, thereby reducing the dependence on high-performance hardware, reducing the complexity of model maintenance, and improving the scalability of the system.

[0114] In summary, the present application effectively solves the problems in the prior art such as limited deployment of large models, insufficient accuracy of small models, lack of intelligence in task scheduling, and high model maintenance costs, and provides a more efficient, intelligent, and flexible model inference solution for automated production robots.

[0115] In one embodiment, as Figure 5 shown, a model generation method is provided, including the following steps:

[0116] Step 701, when the server detects a preset trigger condition, it sends a model update request to the central cloud; the model update request carries the service information of the access network;

[0117] Step 702, the central cloud responds to the model update request, performs knowledge distillation on the pre-deployed large model according to the service information, obtains a small model, returns the small model to the server, and the server receives the small model and deploys the small model in the access network;

[0118] Step 703, the access network obtains a first inference result according to the small model, sends the first inference result to the server, and the server forwards the first inference result to the central cloud;

[0119] Step 704, the central cloud compares the first inference result with the second inference result of the large model, obtains the inference accuracy of the small model, and in the case where the inference accuracy is lower than the preset threshold, returns an optimization signal to the server. When the server receives the optimization signal, it continues to update the small model in the access network.

[0120] In specific implementation, when the server detects a preset trigger condition, it can send a model update request to the central cloud. The model update request carries the service information of the access network. When the central cloud receives the model update request, it can perform knowledge distillation on the pre-deployed large model according to the service information in the model update request, obtain a small model, send the small model to the server, and the server deploys the received small model in the access network. The access network can perform inference on the task data according to the small model, obtain a first inference result, and forward the first inference result to the central cloud through the server. The central cloud compares the first inference result forwarded by the server with the second inference result obtained by its own large model for inferring the task data, obtains the inference accuracy of the small model. If the inference accuracy of the small model is lower than the preset threshold, it means that the inference accuracy of the current small model in the access network fails to meet the expectation. At this time, the central cloud can generate an optimization signal and send the optimization signal to the server. When the server receives the optimization signal, it updates the current small model in the access network, and can either send a model update request to the central cloud again or adjust the parameters of the small model deployed in the access network; otherwise, if the inference accuracy is not lower than the preset threshold, it means that the inference accuracy of the current small model in the access network meets the expectation and there is no need to update. At this time, the current small model in the access network can run independently without the need for the central cloud to cooperate and confirm.

[0121] In the above model generation method, when the server detects a preset trigger condition, it sends a model update request to the central cloud. The central cloud responds to the model update request, determines a small model based on the pre-deployed large model, and the server receives the small model returned by the central cloud and deploys the small model in the access network. When the preset trigger condition is met, the central cloud can be requested to perform knowledge distillation on the pre-deployed large model to dynamically generate a small model adapted to AI RAN for AI RAN to perform inference, reducing the computing requirements of AI RAN and solving the problem of limited computing resources.

[0122] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0123] Based on the same inventive concept, an embodiment of the present application further provides a model generation device for implementing the above-mentioned model generation method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the model generation device provided below can refer to the limitations on the model generation method in the above text, and will not be repeated here.

[0124] In an exemplary embodiment, as Figure 6 shown, a model generation device is provided, which is applied to a server and includes: a sending module 802 and a receiving module 804, where:

[0125] The sending module 802 is configured to send a model update request to the central cloud when a preset trigger condition is detected; the central cloud responds to the model update request and determines a small model based on the pre-deployed large model;

[0126] The receiving module 804 is configured to receive the small model returned by the central cloud and deploy the small model in the access network.

[0127] In an exemplary embodiment, the above model generation device further includes an update module, configured to receive the first inference result of the small model sent by the access network; forward the first inference result to the central cloud; the central cloud compares the first inference result with the second inference result of the large model to obtain the inference accuracy of the small model, and returns an optimization signal when the inference accuracy is lower than a preset threshold; when receiving the optimization signal, update the small model of the access network.

[0128] In an exemplary embodiment, the above update module is further configured to adjust the parameters of the small model to obtain new parameters of the small model; send the new parameters to the access network; the access network updates the small model according to the new parameters.

[0129] In an exemplary embodiment, the above update module is further configured to send the model update request to the central cloud; the central cloud responds to the model update request and determines a new small model according to the large model; deploy the new small model returned by the central cloud in the access network.

[0130] In an exemplary embodiment, the above update module is further configured to receive the inference data uploaded by the access network; when the inference accuracy of the small model is not lower than the preset threshold, the access network uses the small model for inference and statistically analyzes the inference data during the inference process of the small model; forward the inference data to the central cloud; the central cloud updates the large model according to the inference data.

[0131] In an exemplary embodiment, the model update request carries the service information of the access network, and the central cloud performs knowledge distillation on the large model according to the service information to obtain the small model.

[0132] In an exemplary embodiment, there is provided a model generation device applied to a central cloud, including:

[0133] A response module, configured to respond to a model update request from a server and determine a small model according to a pre-deployed large model;

[0134] A return module, configured to return the small model to the server; the server deploys the small model in the access network.

[0135] In an exemplary embodiment, the above response module is further configured to obtain the service information of the access network carried by the model update request; perform knowledge distillation on the large model according to the service information to obtain the small model.

[0136] In an exemplary embodiment, the above model generation device further includes a comparison module, configured to receive the first inference result of the access network forwarded by the server; compare the first inference result with the second inference result of the large model to obtain the inference accuracy of the small model; and return an optimization signal to the server when the inference accuracy is lower than a preset threshold.

[0137] In an exemplary embodiment, there is provided a model generation device, which is applied to an access network and includes

[0138] an acquisition module, configured to acquire a small model; the small model is determined by a cloud based on a pre-deployed large model and is deployed by a server for the determined small model;

[0139] an inference module, configured to infer task data according to the small model to obtain an inference result of the task data.

[0140] Each module in the above model generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in a processor in a communication device in hardware form or be independent of the processor, or can be stored in a memory in the communication device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0141] In an exemplary embodiment, there is provided a communication device, which can be a server, a central cloud, or an access network, and its internal structure diagram can be as Figure 7 shown. The communication device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the communication device is used to provide computing and control capabilities. The memory of the communication device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the communication device is used to store model generation data. The input / output interface of the communication device is used to exchange information between the processor and external devices. The communication interface of the communication device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a model generation method.

[0142] Those skilled in the art can understand that Figure 7The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the communication device to which the solution of this application is applied. The specific communication device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0143] In one embodiment, a communication device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0144] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0145] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0146] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0147] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0148] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in this application.

[0149] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A model generation method, characterized in that, The method is applied to a server and includes: When a preset trigger condition is monitored, sending a model update request to the central cloud; the central cloud responds to the model update request and determines a small model according to a pre-deployed large model; Receiving the small model returned by the central cloud and deploying the small model in the access network.

2. The method according to claim 1, wherein The method further includes: Receiving a first inference result of the small model sent by the access network; Forwarding the first inference result to the central cloud; the central cloud compares the first inference result with a second inference result of the large model to obtain the inference accuracy of the small model, and returns an optimization signal when the inference accuracy is lower than a preset threshold; When receiving the optimization signal, updating the small model in the access network.

3. The method according to claim 2, wherein The updating the small model in the access network includes: Adjusting parameters of the small model to obtain new parameters of the small model; Sending the new parameters to the access network; the access network updates the small model according to the new parameters.

4. The method according to claim 2, wherein The updating the small model in the access network includes: Sending the model update request to the central cloud; the central cloud responds to the model update request and determines a new small model according to the large model; Deploying the new small model returned by the central cloud in the access network.

5. The method according to claim 1, wherein The method further includes: Receiving inference data uploaded by the access network; the access network uses the small model for inference when the inference accuracy of the small model is not lower than a preset threshold, and statistics the inference data during the inference process of the small model; Forwarding the inference data to the central cloud; the central cloud updates the large model according to the inference data.

6. The method according to claim 1, characterized in that, The model update request carries service information of the access network, and the central cloud performs knowledge distillation on the large model according to the service information to obtain the small model.

7. A model generation method, characterized in that, The method is applied to the central cloud and includes: Responding to a model update request of the server and determining a small model according to a pre-deployed large model; Returning the small model to the server; the server deploys the small model in the access network.

8. The method according to claim 7, wherein The responding to the model update request of the server and determining a small model according to a pre-deployed large model includes: Obtaining the service information of the access network carried by the model update request; Performing knowledge distillation on the large model according to the service information to obtain the small model.

9. The method according to claim 7, wherein The method further includes: Receiving the first inference result of the access network forwarded by the server; Comparing the first inference result with a second inference result of the large model to obtain the inference accuracy of the small model; Returning an optimization signal to the server when the inference accuracy is lower than a preset threshold.

10. A model generation method, characterized in that, The method is applied to the access network and includes: Obtaining a small model; the small model is determined by the cloud according to a pre-deployed large model and deployed by the server; Obtaining an inference result of the task data according to the small model inference task data.

11. A model generation device, characterized in that The device is applied to the server and includes: A sending module, configured to send a model update request to the central cloud when a preset trigger condition is detected; the central cloud responds to the model update request and determines a small model according to a pre-deployed large model; A receiving module, configured to receive the small model returned by the central cloud and deploy the small model in the access network.

12. A communication device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

14. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.