Model deployment method and communication device

By receiving the identification information and deployment scope of the target model, a one-time deployment model is realized to multiple entities, solving the problem of multiple deployment models in the prior art, improving deployment efficiency and reducing complexity.

CN120200903APending Publication Date: 2025-06-24HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311777266.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art requires multiple model deployments when deploying models in multiple entities, which are inefficient and have high complexity.

Method used

By receiving the identification information and deployment scope of the target model, the target model is deployed in one go to multiple entities based on this information, reducing deployment complexity and improving efficiency.

Benefits of technology

Shared model deployment among multiple entities is realized, reducing deployment complexity and time, and improving model deployment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200903A_ABST
    Figure CN120200903A_ABST
Patent Text Reader

Abstract

The invention provides a model deployment method and a communication device, relates to the field of communication, and is beneficial to reducing the complexity of deploying a target model for a plurality of entities and improving the model deployment efficiency. The method comprises the following steps: a first device determines and sends identification information of a target model and a deployment range of the target model to a second device, wherein the deployment range of the target model is used for indicating a plurality of entities needing to use the target model for reasoning; correspondingly, the second device receives identification information of the target model and a deployment range of the target model; and the second device deploys the target model based on the identification information of the target model and the deployment range of the target model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and in particular, to a model deployment method and a communication device. Background Art

[0002] In order to improve the intelligence and automation levels of the network, artificial intelligence (AI) and machine learning (ML) technologies are also gradually being applied. For example, there may be a type of service module (such as an inference function) in an access network device that realizes functions through a model. The access network device can realize corresponding functions based on the model to improve the intelligence level of the access network device.

[0003] After the model goes through the training phase and / or the simulation phase, it also needs to go through the deployment phase and be deployed to an entity that undertakes the inference function to execute the inference function on this entity. The model deployment process can be understood as a process of making the model available on an entity that undertakes the inference function. In some examples, one model deployment process can only enable one entity (such as one access network device) to have the ability to obtain the model inference result. However, in some scenarios, there are multiple entities with inference requirements. With the existing methods, multiple model deployment processes need to be performed for multiple entities, resulting in low efficiency. Summary of the Invention

[0004] This application provides a model deployment method and a communication device, which are beneficial to reducing the complexity of deploying a target model for multiple entities and improving the model deployment efficiency.

[0005] In a first aspect, this application provides a model deployment method. This method can be executed by a second device or a component (such as a chip, a chip system, etc.) in the second device. The second device has the ability to provide a model loading service. The second device can be a network management system, a network element management system, a network device, a radio access network device, or any other device or system that can provide a model loading service. This application does not make specific limitations in this regard. The method includes: receiving the identification information of the target model and the deployment scope of the target model, where the deployment scope of the target model is used to indicate multiple entities that need to perform inferences using the target model; and deploying the target model based on the identification information of the target model and the deployment scope of the target model.

[0006] In the embodiments of the present application, the information received by the second device includes the deployment scope of the target model, which is used to indicate multiple entities that need to perform inference using the target model. Based on this deployment scope, the second device can deploy the target model for all or some of the multiple entities indicated by this deployment scope. Compared with some examples, in a scenario where a model needs to be deployed for multiple entities, and the second device needs to receive the identification information of the target model multiple times for multiple entities, the method provided in the present application indicates the deployment scope of the target model during the model deployment process, so that one model deployment process can correspond to multiple entities within this deployment scope that need to perform inference using the target model, reducing the complexity of deploying the target model for multiple entities and facilitating the improvement of the model deployment efficiency.

[0007] In combination with the first aspect, in some implementation manners of the first aspect, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple inference functions; or, identification information of multiple entities.

[0008] It should be understood that the deployment scope of the target model can refer to a scope that includes multiple entities that need to perform inference using the target model, which can be a substantial regional scope or an abstract scope, and the present application does not make any limitations thereto.

[0009] Optionally, the information of the geographical area can be the administrative division information of the geographical area, the longitude and latitude information of the geographical area, etc., and the present application does not make any limitations thereto. It can be understood that there may be a first mapping relationship between the information of the geographical area and the entity undertaking the inference function, and the second device can maintain this first mapping relationship. In one possible implementation manner, the second device can determine, based on the information of the geographical area and the first mapping relationship, multiple entities located in this geographical area as the entities that need to perform inference using the target model. Further, the second device can perform subsequent steps for the determined multiple entities.

[0010] Optionally, the identification information of the network area can be the identification information of any subnet that can distinguish the network area, such as a local area network, a metropolitan area network, a wide area network, etc., and the present application does not make any limitations thereto. It can be understood that there may be a second mapping relationship between the identification information of the network area and the entity undertaking the inference function, and the second device can maintain this second mapping relationship. In one possible implementation manner, the second device can determine, based on the identification information of the network area and the second mapping relationship, multiple entities that need to perform inference using the target model, and further perform subsequent steps for the determined multiple entities.

[0011] In some examples, the entity that undertakes the inference function may be referred to as the inference function (AIMl Inference Function). Optionally, the identification information of multiple inference functions may be used to identify multiple entities that undertake the inference function. Exemplarily, the deployment scope of the target model may be indicated by multiple AIMl Inference Function identifications or a list of AIMl Inference Functions.

[0012] Optionally, the identification information of multiple entities may be any information that can uniquely identify an entity, such as the Internet Protocol (IP) addresses and Media Access Control (MAC) addresses of multiple entities. This application does not make any limitations in this regard. In some examples, the identification information of multiple entities may be presented in the form of a list, but this application does not make specific limitations in this regard.

[0013] In combination with the first aspect, in some implementation manners of the first aspect, deploying the target model based on the identification information of the target model and the deployment scope of the target model includes: sending first information to a target entity, where the first information is used to instruct the target entity to deploy the target model, and the target entity includes all or part of the multiple entities.

[0014] In a possible implementation manner, the target entity is determined by a second device from the deployment scope of the target model, and the target entity is determined by the second device based on the resources and / or capabilities of multiple entities.

[0015] In another possible implementation manner, the target entity is specified by a first device, and the target entity is determined by the first device based on the resources and / or capabilities of multiple entities.

[0016] In combination with the first aspect, in some implementation manners of the first aspect, the target entity includes all of the multiple entities, and the first information includes the identification information of the target model.

[0017] In this embodiment, all of the multiple entities indicated by the deployment scope of the target model are target entities. The manner in which each entity in the multiple entities performs inference using the target model may be: performing inference based on the target model deployed by itself and obtaining an inference result. At this time, the first information sent by the second device to the target entity includes the identification information of the target model to instruct the target entity to deploy the target model.

[0018] In combination with the first aspect, in some implementation manners of the first aspect, the target entity includes part of the multiple entities, and the first information includes the identification information of the target model and the deployment scope of the target model.

[0019] In this embodiment, among the multiple entities indicated by the deployment scope of the target model, there are a target entity and other entities other than the target entity. In such a case, the target entity can perform inference based on the target model deployed by itself, while other entities can perform inference using the target model deployed on the target entity.

[0020] In combination with the first aspect, in some implementation manners of the first aspect, when the target entity is specified by the first device, the method further includes: receiving identification information of the target entity.

[0021] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: sending second information, where the second information is used to indicate that the target model is successfully deployed, and the second information includes the identification information of the target model and the identification information of the target entity.

[0022] It should be understood that after the target model is successfully deployed on the target entity, the second device can notify the first entity (the entity other than the target entity among the multiple entities) of the successful deployment of the target model through the second information. The identification information of the target model in the second information is used to indicate the deployed model, and the identification information of the target entity is used to indicate on which entity the target model is deployed. The deployment scope of the target model can be used by the target entity to authenticate the first entity.

[0023] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: sending first identification information and identification information of a first process, where the first identification information is associated with a message carrying the identification information of the target model, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.

[0024] Optionally, the first identification information and the identification information of the first process can be sent through a response message, and this response message can be a response to a message carrying the identification information of the target model, but this application does not limit this.

[0025] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: receiving first indication information, where the first indication information is used to indicate the management manner of the deployment process of the target model.

[0026] In combination with the first aspect, in some implementation manners of the first aspect, the management manner of the deployment process can be batch management.

[0027] "Batch management" can be understood as regarding the deployment processes of the target models for multiple target entities as a whole. In such a case, the first process corresponds to the overall deployment process of the target model.

[0028] In a possible implementation, the first indication information indicates that the deployment process management method is batch management. In another possible implementation, the deployment process management method of the target model is defaulted to batch management. When the deployment process management method is batch management, the first process corresponds to the overall deployment process of the target model.

[0029] Combined with the first aspect, in some implementations of the first aspect, the second device may report the overall deployment process of the target model.

[0030] In a possible implementation, the second device actively reports the status of the first process. The method further includes: sending third information, where the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process. The status includes: running, canceling, canceled, paused, or ended.

[0031] In another possible implementation, the second device reports the status of the first process in response to a request from the first device. Before the second device sends the third information, the method further includes: the second device receives a second request from the first device, where the second request is used to request the status of the first process, and the first request includes the first identification information, or the identification information of the first process.

[0032] Combined with the first aspect, in some implementations of the first aspect, the method further includes: receiving fourth information, where the fourth information is used to control the overall deployment process of the target model. The fourth information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction. The control instruction includes cancel, pause, or start.

[0033] In the embodiments of the present application, through the model deployment method of batch management, the number of the first processes can be one. The second device can achieve the purpose of deploying the target model to multiple target entities by only maintaining one process instance, which is convenient for the second device to manage the model deployment process. At the same time, in the process of the first device and the second device requesting and reporting the status of the first process and the control interaction of the first process, the process of deploying the target model to multiple target entities is regarded as a whole, without the need to pay attention to the model deployment process of each target entity, nor the need to interact with the information of each target entity, which is beneficial to reducing the signaling overhead between devices and improving the deployment efficiency of the target model.

[0034] Combined with the first aspect, in some implementations of the first aspect, the deployment process management method of the target model may also be independent management.

[0035] "Independent management" can be understood as creating an independent process for management for the model deployment process of each of multiple target entities.

[0036] In one possible implementation, the first indication information indicates that the management method of the deployment process of the target model is default to independent management; in another possible implementation, the deployment process management method is independent management. In the case where the deployment process management method is independent management, the first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of one entity.

[0037] Combined with the first aspect, in some implementations of the first aspect, the second device can report the status of all or part of the multiple processes.

[0038] In one possible implementation, the second device actively reports the status of all or part of the multiple processes, and the method further includes: sending fifth information, where the fifth information includes the first identification information, the identification information of each entity corresponding to all or part of the multiple processes, and the status of each of the all or part of the processes; or, the identification information of the all or part of the processes, and the status of each of the all or part of the processes, and the status includes: running, canceling, canceled, paused, or ended.

[0039] In another possible implementation, the second device reports the status of all or part of the multiple processes to the first device in response to a third request from the first device. Specifically, before the second device sends the fifth information, the second device receives a third request from the first device, and this third request is used to request the status of all or part of the multiple processes. The third request includes the first identification information, the identification information of each entity corresponding to all or part of the multiple processes, or the identification information of all or part of the processes.

[0040] Combined with the first aspect, in some implementations of the first aspect, the method further includes: receiving sixth information, where the sixth information is used to control all or part of the multiple processes, and the sixth information includes the first identification information, the identification information of the entity corresponding to the all or part of the processes, and a control instruction for each of the all or part of the processes; or, the identification information of the all or part of the processes, and the control instruction.

[0041] In the embodiments of the present application, through an independently managed model deployment method, the first process may include multiple processes, and the multiple processes correspond one-to-one to the identification information of multiple target entities. The multiple processes respectively correspond to the target model deployment processes of the respective multiple target entities. In this way, the first device can respectively obtain and control the states corresponding to the model deployment processes of each target entity. In the case where some of the multiple target entities fail to be deployed, it can be determined which target entities have failed to be deployed, and based on the identification information of the target entities that have failed to be deployed, the target models can be redeployed for these target entities, which is beneficial to improving the management accuracy and enhancing the efficiency of model deployment.

[0042] In a second aspect, the present application further provides a model deployment method. This method can be executed by a first device or a component (such as a chip, a chip system, etc.) in the first device. The first device has the ability to call a model loading service. The first device can be a network management system, a network element management system, a network device, or a radio access network device, or any other device or system that has the ability to call a model loading service. The present application does not make specific limitations on this. The method includes: determining the identification information of a target model and the deployment scope of the target model, where the deployment scope of the target model is used to indicate multiple entities that need to perform inference using the target model; sending the identification information of the target model and the deployment scope of the target model.

[0043] In combination with the second aspect, in some implementation manners of the second aspect, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple inference functions; or, identification information of multiple entities.

[0044] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: sending the identification information of a target entity, where the target entity includes some of the multiple entities, and the target entity is determined based on the resources and / or capabilities of the multiple entities.

[0045] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: sending first identification information and the identification information of a first process, where the first identification information is associated with a message carrying the identification information of the target model, and the first identification information has a mapping relationship with the first process, and the first process is used to manage the deployment process of the target model.

[0046] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: receiving first indication information, where the first indication information is used to indicate the management manner of the deployment process of the target model.

[0047] In combination with the second aspect, in some implementation manners of the second aspect, the deployment process management manner is batch management.

[0048] In combination with the second aspect, in some implementation manners of the second aspect, the deployment process management manner of the target model is default to batch management.

[0049] In combination with the second aspect, in some implementation manners of the second aspect, the first process corresponds to the overall deployment process of the target model.

[0050] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: sending third information, where the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.

[0051] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: receiving fourth information, where the fourth information is used to control the overall deployment process of the target model, and the fourth information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause, or start.

[0052] In combination with the second aspect, in some implementation manners of the second aspect, the deployment process management manner of the target model is default to independent management.

[0053] In combination with the second aspect, in some implementation manners of the second aspect, the deployment process management manner is independent management.

[0054] In combination with the second aspect, in some implementation manners of the second aspect, the first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of an entity.

[0055] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: sending fifth information, where the fifth information includes the first identification information, the identification information of each entity corresponding to all or part of the multiple processes, and the status of each of the all or part of the processes; or the identification information of the all or part of the processes, and the status of each of the all or part of the processes, and the status includes: running, canceling, canceled, paused, or ended.

[0056] In combination with the second aspect, in some implementations of the second aspect, the method further includes: receiving sixth information for controlling all or part of the multiple processes, where the sixth information includes the first identification information, the identification information of the entity corresponding to all or part of the processes, and control instructions for each of all or part of the processes; or, the identification information of the entity corresponding to all or part of the processes and the control instructions.

[0057] In combination with the second aspect, in some implementations of the second aspect, the identification information of the target model is sent through a request message or a configuration message.

[0058] In combination with the second aspect, in some implementations of the second aspect, the deployment scope of the target model is sent through a request message or a configuration message.

[0059] In a third aspect, the present application further provides a model deployment method, which can be executed by a third device or a component in the third device (such as a chip, a chip system, etc.). The third device has the ability to undertake the inference function. The third device can be a network management system, a network element management system, a network device or a radio access network device, or other entities, devices or systems that can undertake the inference function. The present application does not make specific limitations on this. The method includes: receiving first information for indicating the deployment of a target model, where the first information includes the deployment scope of the target model and the identification information of the target model, and the deployment scope of the target model is used to indicate multiple entities that need to use the target model for inference; based on the first information, deploying the target model.

[0060] In combination with the third aspect, in some implementations of the third aspect, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple inference functions; or, identification information of multiple entities.

[0061] In combination with the third aspect, in some implementations of the third aspect, the method further includes: receiving a first request for requesting the inference result of the target model, where the first request includes the identification information of the target model and the input parameters of the target model; based on the input parameters, obtaining the inference result through the target model.

[0062] In combination with the third aspect, in some implementations of the third aspect, the method further includes: sending the inference result.

[0063] In combination with a third aspect, in some implementation manners of the third aspect, the method further includes: receiving second information, where the second information is used to indicate that the target model is successfully deployed, and the second information includes the identification information of the target model and the identification information of the target entity.

[0064] In a fourth aspect, the present application provides a communication device, including: a module for executing the method according to any one of the first aspect, the second aspect, or the third aspect.

[0065] In a fifth aspect, the present application provides another communication device, including a processor, which is coupled to a memory and can be used to execute instructions in the memory to implement the method in any possible implementation manner of the first aspect, the second aspect, or the third aspect above. Optionally, the device further includes a memory. Optionally, the device further includes a communication interface, and the processor is coupled to the communication interface.

[0066] In a sixth aspect, the present application provides a communication system, including a first device, a second device, and a third device, where the second device is used to execute the method according to any one of the first aspect, the first device is used to execute the method according to any one of the second aspect, and the third device is used to execute the method according to any one of the third aspect.

[0067] In a seventh aspect, the present application provides a processor, including: an input circuit, an output circuit, and a processing circuit. The processing circuit is used to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method in any possible implementation manner of the first aspect, the second aspect, or the third aspect above.

[0068] In a specific implementation process, the above-mentioned processor may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be transistors, gate circuits, flip-flops, and various logic circuits, etc. The input signal received by the input circuit may be received and input by, for example, but not limited to, a receiver, and the signal output by the output circuit may be output to, for example, but not limited to, a transmitter and transmitted by the transmitter, and the input circuit and the output circuit may be the same circuit, and this circuit is used as the input circuit and the output circuit at different times respectively. The embodiments of the present application do not limit the specific implementation manners of the processor and various circuits.

[0069] In an eighth aspect, a processing device is provided, including a processor and a memory. The processor is used to read instructions stored in the memory and can receive a signal through a receiver and transmit a signal through a transmitter to execute the method in any possible implementation manner of the first aspect, the second aspect, or the third aspect above.

[0070] Optionally, the processor is one or more, and the memory is one or more.

[0071] Optionally, the memory may be integrated with the processor or separately provided from the processor.

[0072] In a specific implementation process, the memory may be a non-transitory memory, such as a read only memory (ROM), which may be integrated with the processor on the same chip or separately provided on different chips. The embodiments of the present application do not limit the type of the memory and the setting manner of the memory and the processor.

[0073] It should be understood that relevant data interaction processes, such as sending indication information, may be a process of outputting indication information from the processor, and receiving capability information may be a process of the processor receiving input capability information. Specifically, the data processed and output may be output to a transmitter, and the input data received by the processor may come from a receiver. Among them, the transmitter and the receiver may be collectively referred to as a transceiver.

[0074] The processing device in the eighth aspect above may be a chip. The processor may be implemented by hardware or by software. When implemented by hardware, the processor may be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor may be a general-purpose processor, which is implemented by reading software code stored in the memory. The memory may be integrated in the processor or may exist independently outside the processor.

[0075] In a ninth aspect, a computer program product is provided. The computer program product includes: a computer program (which may also be referred to as code or instruction). When the computer program is run, the computer is caused to execute the method in any one of the possible implementation manners in the first aspect, the second aspect, or the third aspect above.

[0076] In a tenth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program (which may also be referred to as code or instruction). When it runs on a computer, the computer is caused to execute the method in any one of the possible implementation manners in the first aspect, the second aspect, or the third aspect above. Description of the Drawings

[0077] Figure 1 A model workflow provided by an embodiment of the present application;

[0078] Figure 2 A network resource model working diagram provided by an embodiment of the present application;

[0079] Figure 3 A network resource model working diagram provided by an embodiment of the application;

[0080] Figure 4Schematic diagram of a communication system applicable to the application embodiments;

[0081] Figure 5 Schematic flowchart of a model deployment method provided by the application embodiments;

[0082] Figure 6 Schematic flowchart of a model deployment method provided by the application embodiments;

[0083] Figure 7 Schematic flowchart of another model deployment method provided by the application embodiments;

[0084] Figure 8 Schematic flowchart of yet another model deployment method provided by the application embodiments;

[0085] Figure 9 Schematic flowchart of still another model deployment method provided by the application embodiments;

[0086] Figure 10 Schematic flowchart of a model deployment method provided by the application embodiments;

[0087] Figure 11 Schematic block diagram of a communication device provided by the application embodiments;

[0088] Figure 12 Schematic block diagram of another communication device provided by the application embodiments. Detailed implementation manners

[0089] Next, the technical solutions in the present application will be described with reference to the accompanying drawings.

[0090] Before introducing the present application, the following points are explained first.

[0091] First, in the embodiments shown below, each term and English abbreviation, such as reference data or differential data, etc., are all exemplary examples given for convenience of description, and should not constitute any limitation to the present application. The present application does not exclude the possibility of defining other terms that can achieve the same or similar functions in existing or future protocols.

[0092] Second, in the embodiments shown below, the first, second, and various numerical numbers are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.

[0093] Third, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or a similar expression refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.

[0094] To improve the intelligence and automation levels of the network, artificial intelligence (AI) and machine learning (ML) technologies are also gradually being applied to promote network intelligence.

[0095] Figure 1 An exemplary AI / ML workflow is shown, mainly including a training phase, an emulation phase, a deployment phase, and an inference phase. It should be understood that generally, the AI / ML workflow is implemented sequentially according to the training phase, emulation phase, deployment phase, and inference phase. However, in some possible implementation manners, the AI / ML model can execute the inference phase after the training phase and / or the emulation phase. Or, after the training phase, the emulation phase can be skipped and the deployment phase and inference phase can be executed. Or, after the emulation phase or the inference phase, it can enter the training phase again. Among them, the deployment phase is the process of loading the model into the inference function. In a possible implementation manner, after the model goes through the training phase and the emulation phase, it is loaded into the corresponding inference function through the deployment phase to enter the inference phase and execute the inference process.

[0096] In some examples, the role of providing the ML entity loading management service (MnS) is called the producer, and the role of invoking the ML entity loading management service is called the consumer. Two sub-use cases are defined for the deployment phase and solutions for these two sub-use cases.

[0097] The two sub-use cases include:

[0098] Sub - use case 1: The consumer requests the producer to perform model loading (consumer requested ML entityloading).

[0099] Sub - use case 2: The consumer configures a policy for the producer, and the producer triggers model loading by itself (control of producer - initiated ML entity loading).

[0100] Solutions for the two sub - use cases:

[0101] The current solution defines 3 information object classes (IOCs), including: MLEntityLoadingRequest, MLEntityLoadingPolicy, and MLEntityLoadingProcess for representing the ML model loading process. Among them, MLEntityLoadingRequest and MLEntityLoadingProcess are applied to the scenario of the above sub - use case 1, and MLEntityLoadingPolicy and MLEntityLoadingProcess are applied to the scenario of the above sub - use case 2.

[0102] The attributes of the above 3 information object classes are described in detail below.

[0103] 1. MLEntityLoadingRequest< <ioc>>

[0104] MLEntityLoadingRequest< <ioc>>Represents an ML model loading request created by a consumer. The consumer uses this IOC request to ask the producer to load the ML model into the target inference function. The attributes of this IOC are shown in Table 1.

[0105] Table 1

[0106]

[0107] Among them, the requestStatus attribute represents the status of the request. The value of the requestStatus attribute can be any of the following: Not Started (NOT_STARTED), Loading in Progress (LOADING_IN_PROGRES), Suspended (SUSPENDED), Finished Successfully (FINISHED_SUCCESS), Finished with Failure (FINISHED_FAILED), or Cancelled (CANCELLED). The cancelRequest attribute represents whether the consumer cancels the ML entity loading request, with values of Yes (TRUE) or No (FALSE). The role-related attribute mLEntityToLoadRef represents the identifier of the model to be loaded for the ML model loading request created by the consumer.

[0108] It should be understood that the values of the "support qualifier" include mandatory (M), optional (O), conditionally optional (CO), and conditionally mandatory (CM). The value M indicates that the attribute is a mandatory attribute, the value O indicates that the attribute is an optional attribute, the value CM indicates conditionally mandatory, that is, the attribute is a mandatory attribute when a certain condition is met, and the value CO indicates that the attribute is an optional attribute when a certain condition is met; the values of "readable or not" include true (TRUE, T) and false (FALSE, F). When the value is T, it means the attribute is readable, and when the value is F, it means the attribute is not readable; the values of "writable or not" also include TRUE and FALSE. When the value is T, it means the attribute is writable, and when the value is F, it means the attribute is not writable. This explanation also applies to the following tables and will not be repeated later.

[0109] 2. MLEntityLoadingPolicy< <ioc>>

[0110] MLEntityLoadingPolicy< <ioc>>Indicates the ML model loading strategy set by the consumer for the producer. The consumer uses this IOC to set the conditions for the producer to trigger the ML model loading. Only when the ML model meets the conditions can the producer trigger the model loading. The attributes of this IOC are shown in Table 2.

[0111] Table 2

[0112]

[0113] Among them, the inferenceType attribute is used to indicate the inference type of the initial model associated with the ML model loading strategy set by the consumer for the producer. This attribute is a conditionally required attribute, that is, when the model associated with the ML model loading strategy set by the consumer for the producer is an initial training model, this attribute is a required attribute. It should be understood that the initial training model is the model trained for the first time, and the initial training model has not been assigned a model identifier. Therefore, the inference type of the initial training model can be used to identify the model.

[0114] The mLEntityId attribute is used to indicate the identifier of the retrained model corresponding to the ML model loading strategy set by the consumer for the producer. This attribute is a conditionally required attribute, that is, when the model associated with the ML model loading strategy set by the consumer for the producer is a retrained model, this attribute is a required attribute. It should be understood that an mLEntityId can uniquely identify a model that has undergone initial training.

[0115] The policyForLoading attribute is used to indicate the ML model loading strategy set by the consumer for the producer. This strategy can be a series of thresholds. For example, this strategy can be that the accuracy of the ML model is greater than 90%, or it reaches a preset time point or preset period, or a certain network performance index is lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule).

[0116] 3. MLEntityLoadingProcess< <ioc>>

[0117] MLEntityLoadingProcess< <ioc>>Represents the ML model loading process. In the scenario of the above Sub-use Case 1, the producer can use this IOC to instantiate one or more ML model loading processes for each ML model loading request of the consumer; in the scenario of the above Sub-use Case 2, the producer can use this IOC to instantiate one or more ML model loading processes and associate these one or more ML model loading processes with the ML model loading policy set by the consumer for the producer. The attributes of this IOC are shown in Table 3.

[0118] Table 3

[0119]

[0120] Among them, the progressStatus attribute represents the status of the ML model loading process, and the value of this attribute can be any one of the following: Running, Cancelling, Suspended, Finished, or Cancelled. The cancelProcess attribute represents whether the consumer cancels the ML entity loading process, and the value is TRUE or FALSE. The suspendProcess attribute represents whether the consumer suspends the ML entity loading process, and the value is TRUE or FALSE. The resumeProcess attribute represents whether the consumer resumes the ML entity loading process, and the value is TRUE or FALSE.

[0121] The role-related attribute MLEntityLoadingRequestRef is the loading request identifier associated with the ML entity loading process and is a required attribute when the ML entity loading process corresponds to the scenario of the above Sub-use Case 1. MLEntityLoadingPolicyRef is the loading policy identifier associated with the ML entity loading process and is a required attribute when the ML entity loading process corresponds to the scenario of the above Sub-use Case 2. LoadedMLEntityRef is the model identifier associated with the ML entity loading process.

[0122] Figure 2 Exemplarily shows the network resource model (NRM) relationship diagrams of the respective IOCs corresponding to Tables 1, 2, and 3 above. As Figure 2 As shown, the AiMlInferenceFunction represents a function that can perform the ML inference function and is also an IOC. In the figure, the solid diamonds on the side of the AiMlInferenceFunction object for the connecting lines between the MLEntityLoadingRequest object, MLEntityLoadingProcess object, and MLEntityLoadingPolicy object and the AiMlInferenceFunction object are used to represent the inclusion relationship, that is, the AiMlInferenceFunction object includes the MLEntityLoadingRequest object, MLEntityLoadingProcess object, and MLEntityLoadingPolicy object. Further, the connecting lines between the MLEntityLoadingRequest object and MLEntityLoadingProcess object and the MLEntity object are arrows on the side of the MLEntity object. These arrows are used to represent the association relationship, that is, the MLEntityLoadingRequest object and MLEntityLoadingProcess object are respectively associated with the MLEntity object. Therefore, it can be further understood that the AiMlInferenceFunction object includes the MLEntity object, which is represented in the figure as a solid diamond on the side of the AiMlInferenceFunction object for the connecting line between the AiMlInferenceFunction object and the MLEntity object.

[0123] In addition, the numbers and symbols on the connecting lines between the objects in the figure represent the cardinality relationships between the intended object attributes with inclusion and association relationships. For example Figure 2 On the connection line between the MLEntityLoadingRequest object and the AiMlInferenceFunction object shown in [figure], the cardinality on the MLEntityLoadingRequest object side is "*", and the cardinality on the AiMlInferenceFunction object side is "1". On the connection line between the MLEntityLoadingRequest object and the MLEntity object, the cardinality on both sides of the MLEntityLoadingRequest object and the MLEntity object is "1", which means that an AiMlInferenceFunction object can contain multiple MLEntityLoadingRequest objects, and each MLEntityLoadingRequest object can only be associated with one MLEntity object. Figure 2 The connection rules of other connection lines in [figure] and the subsequent NRM diagrams can also be explained according to this rule, so no more details will be given.

[0124] From Figure 2 it can be seen that the MLEntityLoadingRequest object, the MLEntityLoadingProcess object, and the MLEntityLoadingPolicy object are included under the AiMlInferenceFunction object. For consumers, a MLEntityLoadingRequest or a MLEntityLoadingPolicy created during the model deployment process is for one inference function (such as a software or entity undertaking the inference function), that is, during the model deployment process, the target inference function explicitly or implicitly indicated by the MLEntityLoadingRequest or MLEntityLoadingPolicy is only one. Therefore, the producer can only deploy the model for the indicated target inference function. If there is a situation where models need to be deployed for multiple inference functions, then the consumer needs to maintain a network resource model relationship as shown in Figure 2 for each inference function, that is, a MLEntityLoadingRequest or a MLEntityLoadingPolicy needs to be created for each inference function to achieve model deployment, which is a complex process and has low efficiency.

[0125] In view of this, the present application provides a model deployment method and a communication device, which indicate the deployment scope of a target model during the model deployment process, so that a single model deployment process can correspond to multiple entities that need to perform inference using the target model within the deployment scope, reducing the complexity of deploying the target model for multiple entities and improving the model deployment efficiency.

[0126] It should be understood that in the embodiments of the present application, the inference function (AiMlInferenceFunction) can be understood as an entity that undertakes the inference function, such as a base station. To avoid ambiguity, in the following description, "entity" is used to represent the inference function and will not be elaborated further.

[0127] Next, a detailed description of the design for IOC proposed in the present application will be given first.

[0128] 1. Add the following attributes to the model loading request object created by the consumer:

[0129] Attribute 1: Used to indicate the deployment scope of the target model corresponding to the model loading request created by the consumer, and this deployment scope can be used to indicate multiple entities that need to perform inference using the target model;

[0130] Attribute 2: Used to indicate the management method during the deployment process of the target model, for example, it can be batch management or independent management;

[0131] Attribute 3: Used to indicate the entity that the consumer specifically designates to actually deploy the target model.

[0132] Exemplarily, the model loading request created by the consumer can be through MLEntityLoadingRequest< <ioc>>indicates that Table 4 exemplarily shows an MLEntityLoadingRequest provided by an embodiment of the present application< <ioc>>Contained attributes.

[0133] Table 4

[0134]

[0135]

[0136] In a possible implementation, the above-mentioned attribute 1 can be the role-related attribute aIMlInferenceFunctionRef. This attribute can be a conditional optional attribute. For example, when the model loading request created by the consumer targets entities within a certain range, this attribute is optional. For example, when the deployment range is a geographical area, optionally, a specific list of inference functions within the geographical area can be further specified to select a part of the specified inference functions within a geographical area range, or it can be not specified, in which case all inference functions within a geographical area range will be selected; the above-mentioned attribute 2 can be infFuncBatchIndicator. This attribute is an optional attribute, and the management method during the model deployment process can be indicated by the value of infFuncBatchIndicator. For example, batch management can be performed on the model deployment process for multiple entities, or process-independent management can be performed on the model deployment of each entity. It should be understood that the embodiments of the present application do not make specific limitations on the manifestation forms of attribute 1 and attribute 2.

[0137] In a possible implementation, the above-mentioned attribute 3 can be the role-related attribute aIMlInferenceFunctionToLoadMlEntityRef. This attribute can be a conditional optional attribute. For example, in the case where the model loading request created by the consumer is used to indicate the deployment of the target model for some inference functions within the above-mentioned deployment range, this attribute is optional. The value of this attribute can represent the identifier of the part of the inference functions specified by the consumer, but the present application does not make specific limitations on this.

[0138] 2. Add the following attributes to the model loading policy object set by the consumer for the producer:

[0139] Attribute 4: Used to indicate the deployment range of the target model corresponding to the model loading policy set by the consumer for the producer. This deployment range can be used to indicate multiple entities that need to perform inferences using the target model;

[0140] Attribute 5: Used to indicate the entities that the consumer specifically needs to deploy the target model to.

[0141] Exemplarily, the model loading policy set by the consumer for the producer can be through MLEntityLoadingPolicy< <ioc>>Indicates that Table 5 exemplarily shows an MLEntityLoadingPolicy provided by an embodiment of the present application< <ioc>>Included attributes.

[0142] Table 5

[0143]

[0144] In one possible implementation, the above-mentioned Attribute 4 may be the role-related attribute aIMlInferenceFunctionRef, which can be a conditional optional attribute. For example, when the model loading strategy set by the consumer for the producer targets entities within a certain range, this attribute is optional. It should be understood that Attribute 4 may also have other forms, which are not limited in this application.

[0145] In one possible implementation, the above-mentioned Attribute 5 may be IMlInferenceFunctionToLoadMlEntityRef, which can be a conditional optional attribute. For example, it may be optional when the model loading strategy set by the consumer for the producer includes deploying the target model for some inference functions within the deployment scope of the target model. The value of this attribute can represent the identifier of the part of the inference function specified by the consumer, but this application does not make any limitations in this regard.

[0146] In one possible implementation, the model loading process object can be passed through MLEntityLoadingProcess< <ioc>>It is indicated that the attributes it contains can be as shown in Table 3, but this application does not limit this.

[0147] Figure 3 Exemplarily shows a network resource model relationship diagram provided by an embodiment of this application. Compared with Figure 2 the MLEntityLoadingRequest object, the MLEntityLoadingProcess object, and the MLEntityLoadingPolicy object are no longer included in an AiMlInferenceFunction object, but are included in a ManagedEntity object, and this ManagedEntity object is a proxy class (ProxyClass).

[0148] Exemplarily, the ManagedEntity can represent one or more of the following: a group of managed entities, which can be represented by a network area, such as a subnetwork (SubNetwork), and can include a group of base stations; a group of managed functions (ManagedFunction), such as a software-implemented communication function running on dedicated hardware, or a communication function represented by software running on a network function virtualization infrastructure (NFVI), which can be a virtual network function of a base station or a core network; or, a group of managed management functions (ManagementFunction), such as an inference function (AI / ML inference function), which can also be understood as the bearer of the inference function.

[0149] In addition, on the arrow line where the MLEntityLoadingProcess object points to the MLEntityLoadingPolicy object, and on the arrow line where the MLEntityLoadingProcess object points to the MLEntityLoadingRequest object, the cardinality on the side close to the MLEntityLoadingProcess object is modified from "*" ( Figure 2 shown in ) to "1..*", indicating that the number of MLEntityLoadingProcess objects associated with one MLEntityLoadingPolicy object can be at least one, and the number of MLEntityLoadingProcess objects associated with one MLEntityLoadingRequest object can also be at least one.

[0150] It should be understood that based on Figure 3 In the network resource model shown, an MLEntityLoadingRequest or an MLEntityLoadingPolicy created by a consumer during a single model deployment process can no longer be targeted at only one entity (such as Figure 2 an AiMlInferenceFunction shown), but can be targeted at multiple entities (such as Figure 3 a ManagedEntity shown).

[0151] Figure 4 FIG. is a schematic diagram of a communication system 400 applicable to an embodiment of the present application. As Figure 4 shown, the communication system 400 includes a first device 410, a second device 420, and a third device 430. Communication can be established between the first device 410, the second device 420, and the third device 430.

[0152] In a possible implementation, the role of the first device 410 can be that of a model loading consumer (MLloading consumer) that can call a model loading service. The role of the second device 420 can be that of a model loading producer (ML loading producer) that can provide a model loading service. The third device 430 can be a role that implements an inference function based on the model (which can also be understood as the bearer of the inference function, the executor of the inference function, or the acquirer of the inference result. In some implementations, this role can be directly referred to as the inference function), but the present application does not make specific limitations in this regard.

[0153] It should be understood that Figure 4 in the communication system shown, the number of the first device and the second device can each be one or more, and the third device can be multiple. The present application does not make limitations in this regard.

[0154] In a possible scenario, the first device 410 can call the model loading service provided by the second device 420 to instruct the target model to be loaded into the specified third device 430, and the third device 430 is responsible for inputting data into the model to obtain the corresponding output.

[0155] In a possible implementation, the first device 410 (or the second device 420, or the third device 430) can include one or more of the following: a module that can implement the functions corresponding to the model loading consumer, a module that can implement the functions corresponding to the model loading producer, or a module that can implement an inference function based on the model. The present application does not make specific limitations in this regard.

[0156] Optionally, the first device 410 may be a network management system (NMS), an element management system (EMS), or a radio access network (RAN) device, or any other device or system that can act as a model loading consumer. This application does not make specific limitations thereon.

[0157] Optionally, the second device 420 may be a network management system, an element management system, or a radio access network device, or any other device or system that can act as a model loading producer. This application does not make specific limitations thereon.

[0158] Optionally, the third device 430 may be a network management system, an element management system, or a radio access network device, or any other device or system that can implement an inference function based on a model. This application does not make specific limitations thereon.

[0159] Optionally, the above-mentioned element management system can manage the network, be responsible for the operation, management, and function maintenance of the network, and may also be referred to as a cross-domain management system. This application does not make limitations thereon.

[0160] Optionally, the element management system can be used to manage one or more network elements of a certain category, and may also be referred to as a domain management system or a single-domain management system. This application does not make limitations thereon.

[0161] Optionally, the above access network device may be a transmission reception point (TRP), or may be an evolved NodeB (eNB or eNodeB) in a Long Term Evolution (LTE) system, or may be a home base station (e.g., home evolved NodeB, or home Node B, HNB), a base band unit (BBU), or may be a radio controller in a Cloud Radio Access Network (CRAN) scenario, or the access network device may be a relay station, an access point, a vehicle-mounted device, a wearable device, and an access network device in a 5G network or an access network device in a future evolved Public Land Mobile Network (PLMN) network, etc. It may be an access point (AP) in a Wireless Local Area Network (WLAN), may be a gNB in a New Radio (NR) system, may be a satellite base station in a satellite communication system, etc. The embodiments of the present application do not limit this.

[0162] Optionally, the above access network device may include a centralized unit (CU) node, or a distributed unit (DU) node, or an access network device including a CU node and a DU node, or an access network device including a control plane CU node (CU-CP node), a user plane CU node (CU-UP node), and a DU node. Among them, the access network device including a CU node and a DU node can split the protocol layer of the access network device. The functions of some protocol layers are centrally controlled by the CU, and the functions of the remaining part or all protocol layers are distributed in the DU, and the DU is centrally controlled by the CU. As an implementation, the protocol stack deployed by the CU includes a radio resource control (RRC) layer, a packet data convergence protocol (PDCP) layer, and a service data adaptation protocol (SDAP) layer. The protocol stack deployed by the DU includes a radio link control (RLC) layer, a media access control (MAC) layer, and a physical layer (PHY) layer. Thus, the CU has the processing capabilities of RRC, PDCP, and SDAP. The DU has the processing capabilities of RLC, MAC, and PHY. The above function splitting is only an example and does not constitute a limitation on the CU and DU. That is to say, there can be other ways of function splitting between the CU and the DU, which are not elaborated in the embodiments of the present application. The functions of the CU can be implemented by one entity or by different entities. For example, the functions of the CU can be further split. For example, the control plane (CP) and the user plane (UP) are separated, that is, the control plane (CU-CP) and the CU user plane (CU-UP) of the CU. For example, the CU-CP and the CU-UP can be implemented by different functional entities. The CU-CP and the CU-UP can be coupled with the DU to jointly complete the functions of the access network device. In a possible way, the CU-CP is responsible for the control plane functions, mainly including RRC and PDCP-C, where PDCP-C is mainly responsible for encryption, decryption, integrity protection, data transmission, etc. of the control plane data. The CU-UP is responsible for the user plane functions, mainly including SDAP and PDCP-U, where SDAP is mainly responsible for processing the data of the core network device and mapping the data flow to the bearer. PDCP-U is mainly responsible for encryption, decryption, integrity protection, header compression, sequence number maintenance, data transmission, etc. of the data plane. Among them, the CU-CP and the CU-UP are connected through the E1 interface. The CU-CP represents that the access network device is connected to the core network device through the interface between the core network device and the access network device. It is connected to the DU through F1-C (control plane).The CU-UP is connected to the DU via F1-U (user plane). In addition, there is also a possible implementation where PDCP-C is also in the CU-UP, which is not limited in this application.

[0163] To make the objectives and technical solutions of this application clearer and more intuitive, the model deployment method and communication device provided in the embodiments of this application will be described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0164] Below, first in conjunction with Figures 5 to 10 A model deployment method provided in the embodiments of this application will be described in detail.

[0165] The model deployment method provided in the embodiments of this application can be implemented based on the design provided in this application for the IOC. This method can be applied to a communication system 400 as shown in Figure 4 or other communication systems, which is not limited in this application. Exemplarily, in the embodiments of this application, taking the role of the first device as the consumer and the role of the second device as the producer as an example, the model deployment method provided in this application will be described from the perspective of interaction between devices.

[0166] Figure 5 FIG. 500 is a schematic flowchart of a model deployment method 500 provided in the embodiments of this application. The method 500 includes the following steps:

[0167] S501. The first device determines the identification information of the target model and the deployment scope of the target model, and the deployment scope of the target model is used to indicate a plurality of entities that need to perform inference using the target model.

[0168] It should be noted that the model in this application can be a machine learning model (ML model) or can be embodied in the form of an entity, such as a machine learning entity (ML entity), which is not limited in this application.

[0169] S502. The first device sends the identification information of the target model and the deployment scope of the target model. Correspondingly, the second device receives the identification information of the target model and the deployment scope of the target model.

[0170] S503. The second device deploys the target model based on the identification information of the target model and the deployment scope of the target model.

[0171] It should be noted that the second device deploying the target model means that the second device loads the target model into one or more target inference functions corresponding to the deployment scope, where the target model is used to perform inference.

[0172] In the embodiments of the present application, the information sent by the first device to the second device includes the deployment scope of the target model, which is used to indicate multiple entities that need to perform inference using the target model. Based on this deployment scope sent by the first device, the second device can deploy the target model for all or part of the multiple entities indicated by this deployment scope. Compared with some examples where, in a scenario where a model needs to be deployed for multiple entities, the first device needs to send the identification information of the target model to the second device multiple times for multiple entities, the method provided by the present application indicates the deployment scope of the target model during the model deployment process, enabling a single model deployment process to correspond to multiple entities within this deployment scope that need to perform inference using the target model, reducing the complexity of deploying the target model for multiple entities and improving the model deployment efficiency.

[0173] As an optional embodiment, the deployment scope of the target model in S501 above can be indicated by one or more of the following: information of a geographical region; identification information of a network region; identification information of multiple inference functions; or, identification information of multiple entities.

[0174] It should be understood that the deployment scope of the target model can refer to a scope that includes multiple entities that need to perform inference using the target model, which can be a substantial regional scope or an abstract scope, and the present application does not limit this.

[0175] Optionally, the information of the geographical region can be administrative division information of the geographical region, longitude and latitude information of the geographical region, etc., and the present application does not limit this. It can be understood that there can be a first mapping relationship between the information of the geographical region and the entity undertaking the inference function, and the first device and / or the second device can maintain this first mapping relationship. In one possible implementation manner, the second device can, based on the information of the geographical region indicated by the first device and the first mapping relationship, determine multiple entities in this geographical region as the entities that need to perform inference using the target model. Further, the second device can perform subsequent steps for the determined multiple entities.

[0176] Optionally, the identification information of the network region can be identification information of any subnet that can distinguish network regions, such as a local area network, a metropolitan area network, a wide area network, etc., and the present application does not limit this. It can be understood that there can be a second mapping relationship between the identification information of the network region and the entity undertaking the inference function, and the first device and / or the second device can maintain this second mapping relationship. In one possible implementation manner, the second device can, based on the identification information of the network region indicated by the first device and the second mapping relationship, determine multiple entities that need to perform inference using the target model, and further perform subsequent steps for the determined multiple entities.

[0177] In some examples, the entity undertaking the inference function may be referred to as the AIMl Inference Function. Optionally, the identification information of multiple inference functions may be used to identify multiple entities undertaking the inference function. Exemplarily, the deployment scope of the target model may be indicated by multiple AIMl Inference Function identifications or a list of AIMl Inference Functions.

[0178] Optionally, the identification information of multiple entities may be any information that can uniquely identify an entity, such as the internet protocol (IP) addresses or media access control (MAC) addresses of multiple entities. This application does not make any limitations in this regard. In some examples, the identification information of multiple entities may be presented in the form of a list, but this application does not make specific limitations in this regard.

[0179] As an alternative embodiment, a possible implementation manner of S503 includes: the second device sends first information to the target entity, and the first information is used to instruct the target entity to deploy the target model.

[0180] It should be understood that the deployment scope indication of the target model refers to multiple entities that need to perform inferences using the target model. In the embodiments of this application, the target entity may be understood as the entity that actually deploys the target model, and the target entity may include all or part of the multiple entities indicated by the deployment scope of the target model.

[0181] In a possible implementation manner, the target entity is determined by the second device from the deployment scope of the target model, and the target entity is determined by the second device based on the resources and / or capabilities of multiple entities.

[0182] In another possible implementation manner, the target entity is specified by the first device. Then, before the second device sends the first information to the target entity, the method further includes: the first device sends the identification information of the target entity to the second device, and the target entity is determined by the first device based on the resources and / or capabilities of multiple entities. Correspondingly, the second device receives the identification information of the target entity.

[0183] In a possible scenario, the target entity includes all entities among multiple entities, and the first information includes the identification information of the target model.

[0184] In this scenario, multiple entities indicated by the deployment scope of the target model are all target entities. The way each entity among the multiple entities performs inference using the target model can be: performing inference based on the target model deployed on itself and obtaining an inference result. At this time, the first information sent by the second device to the target entity includes the identification information of the target model to indicate that the target entity deploys the target model.

[0185] In some examples, the target model has been stored in the target entity. The second device sends the first information to the target entity, and the first information includes the identification information of the target model to indicate that the target entity deploys the target model. The second device's instruction for the target entity to deploy the target model can be understood as the second device instructing the target entity to activate the target model.

[0186] In other examples, the target model is not stored in the target entity. Then, when the second device sends the first information to the target entity to indicate that the target entity deploys the target model, it can be understood that the second device instructs the target entity to download and activate the target model. In some implementations, the first information further includes the storage address of the target model, and the target entity downloads the target model based on this storage address; in other implementations, the target entity maintains a correspondence between the identification information of the target model and the storage address of the target model. The target entity can search for and obtain this storage address based on the identification information of the target model included in the first information, and download the target model based on this storage address.

[0187] In another possible scenario, the target entity includes some entities among the multiple entities, and the first information includes the identification information of the target model and the deployment scope of the target model.

[0188] In this scenario, among the multiple entities indicated by the deployment scope of the target model, there are target entities and other entities other than the target entities. In this case, the target entity can perform inference using the target model deployed on itself, while other entities can perform inference using the target model deployed on the target entity.

[0189] In a possible implementation manner, taking the other entity as the first entity as an example, the ways in which the other entity can perform inference using the target model deployed on the target entity include: the first entity sends a first request to the target entity, and the first request is used to request the inference result of the target model. The first request includes the identification information of the target model and the input parameters of the target model; correspondingly, the target entity receives the first request, and when determining that the first entity belongs to the deployment scope of the target model, based on this input parameter, obtains the inference result through the target model deployed on the target entity. Further, the target entity sends this inference result to the first entity, and correspondingly, the first entity receives this inference result.

[0190] In a possible implementation, before the first entity sends a first request to the target entity, it further includes: a second device sends second information to the first entity, and the second information is used to indicate that the target model is successfully deployed. The second information includes the identification information of the target model and the identification information of the target entity. Correspondingly, the first entity receives the second information.

[0191] It should be understood that after the target model is successfully deployed on the target entity, the second device can notify the first entity that the target model is successfully deployed through the second information. The identification information of the target model in the second information is used to indicate the successfully deployed model, and the identification information of the target entity is used to indicate on which entity the target model is deployed.

[0192] As an optional embodiment, after S502 above, method 500 further includes: the second device sends the first identification information and the identification information of the first process. The first identification information is associated with the message carrying the identification information of the target model, and there is a mapping relationship between the first identification information and the first process. The first process is used to manage the deployment process of the target model. Correspondingly, the first device receives the first identification information and the identification information of the first process.

[0193] Optionally, the first identification information and the identification information of the first process can be sent through a response message, and the response message can be a response to the message carrying the identification information of the target model, but this application does not limit this.

[0194] It should be understood that the embodiments of this application are applicable to the scenarios corresponding to the two sub-use cases defined in the model deployment phase mentioned above. In different scenarios, the identification information of the target model can be carried by different types of messages, so the meaning of the first identification information is different in different scenarios.

[0195] Scenario 1: The scenario corresponding to sub-use case 1 mentioned above. In a possible implementation of this scenario, the identification information of the target model in S501 above is sent by the first device through a request message.

[0196] In a possible implementation, the request message is used to request the second device to deploy the target model for multiple entities indicated by the deployment scope of the target model. That is, it can correspond to the situation where the target entity includes all entities among multiple entities described above.

[0197] In another possible implementation, the request message is used to request the second device to deploy the target model for some of the multiple entities indicated by the deployment scope of the target model. That is, it can correspond to the situation where the target entity includes some entities among multiple entities described above.

[0198] Optionally, the deployment scope of the target model sent by the first device to the second device in S501 above can also be sent through a request message. The deployment scope of the target model and the identification information of the target model can be sent through the same request message or through different request messages. This application does not make any limitations in this regard.

[0199] It should be understood that in this scenario, the first device determines the timing of sending the above request message based on the deployment policy of the target model. The deployment policy can be the requirement for the inference ability of the target model, or the parameters of each functional entity involved in the model deployment process, etc. The deployment policy can be determined by the first device or reported by the entity that needs to perform inference using the target model. This application does not make any limitations in this regard. Exemplarily, the deployment policy can be that the accuracy of the target model is greater than 90%, or it reaches the preset time point or preset period for deploying the target model in terms of time, or it is detected that a certain network performance indicator of the entity is lower than the preset threshold (such as: energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations in this regard.

[0200] In a possible implementation manner, when the first device determines that the target model has met the conditions corresponding to the deployment policy, and / or all or part of the entities indicated by the deployment scope of the target model have met the deployment conditions, the first device sends a request message to the second device. The request message is used to request the deployment of the target model, and the request message includes the identification information of the target model and the deployment scope of the target model. Correspondingly, the second device receives the request message and creates an instance of the request message and a first process instance.

[0201] It should be understood that the first process is used to manage the deployment process of the target model. The first process can be called the deployment process of the target model or the loading process of the target model. This application does not make specific limitations on the specific name of the first process. In addition, the "loading" and "deployment" described in the embodiments of this application can be understood to have the same meaning without special explanation, and will not be repeated later.

[0202] Exemplarily, in this scenario, the request message can correspond to the MLEntityLoadingRequest object designed in this application, and its attributes can be as shown in Table 4.

[0203] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingRequest object, and the identification information of the target model corresponding to the request message can be represented by the value of the role-related attribute mLEntityToLoadRef of the MLEntityLoadingRequest object. However, the present application does not make specific limitations on this.

[0204] Correspondingly, the first process can be a process created based on the MLEntityLoadingProcess object shown in Table 3. The request message can be represented by the value of the role-related attribute MLEntityLoadingRequestRef of the MLEntityLoadingProcess object, and the identification information of the target model associated with the first process can be represented by the value of the role-related attribute LoadedMLEntityRef.

[0205] In the above Scenario 1, an instance of the request message can be an MLEntityLoadingRequest instance, the first identification information can be the identification of the MLEntityLoadingRequest instance, an instance of the first process can be an MLEntityLoadingProcess instance, and the identification information of the first process can be the identification of the MLEntityLoadingProcess instance.

[0206] Scenario 2: The scenario corresponding to Sub-use case 2 mentioned above. In a possible implementation of this scenario, the identification information of the target model in S501 above is sent by the first device through a configuration message, and the configuration message also includes the deployment policy of the target model.

[0207] In a possible implementation, the first device sends the identification information of the target model and the deployment policy of the target model to the second device through a configuration message. After receiving the identification information of the target model and the deployment policy of the target model, the second device creates an instance of the configuration message, and creates an instance of the first process after determining that the target model has met the conditions corresponding to the deployment policy, and / or all or part of the entities indicated by the deployment scope of the target model have met the deployment conditions.

[0208] Optionally, the deployment scope of the target model sent by the first device to the second device in S501 above can also be sent through a configuration message. The deployment scope of the target model, the identification information of the target model, and the deployment policy of the target model can be sent through the same configuration message or multiple configuration messages. The present application does not limit the number, content, and order of the configuration messages.

[0209] Optionally, the deployment policy of the target model in this scenario can be that the accuracy of the target model is greater than 90%, or the time reaches the preset time point or preset period for deploying the target model, or a certain network performance metric of the third device is detected to be lower than the preset threshold (e.g., energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations on this.

[0210] Exemplarily, in this scenario, the configuration message can correspond to the MLEntityLoadingPolicy object designed in this application, and its attributes can be as shown in Table 5.

[0211] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingPolicy object, the deployment policy of the target model can be represented by the value of the attribute policyForLoading, and the identification information of the target model can be represented by the value of the attribute inferenceType or the value of the attribute mLEntityId, but this application does not make limitations on this.

[0212] Correspondingly, the first process can be a process created based on the MLEntityLoadingProcess object described in Table 3. The configuration message can be identified by the value of the role-related attribute MLEntityLoadingPolicyRef of the MLEntityLoadingProcess object, and the identification information of the target model associated with the first process can be represented by the value of the role-related attribute LoadedMLEntityRef.

[0213] In the above Scenario 2, the instance of the configuration message can be an MLEntityLoadingPolicy instance, the first identification information can be the identification of the MLEntityLoadingPolicy instance, the instance of the first process can be an MLEntityLoadingProcess instance, and the identification information of the first process can be the identification of the MLEntityLoadingProcess instance.

[0214] The first process in the above two scenarios is used to manage the deployment process of the target model. When the number of target entities that need to deploy the target model is multiple, the process management methods for deploying the target model to the target entities can include two methods: "batch management" and "independent management", but this application does not make specific limitations on this. Under different deployment process management methods, the meaning of the first process is different. The following takes "batch management" and "independent management" as examples for detailed description.

[0215] The first method: batch management

[0216] "Batch management" can be understood as regarding the deployment process of the target model for multiple target entities as a whole. In this case, the first process corresponds to the overall deployment process of the target model. Exemplarily, the second device creates an MLEntityLoadingProcess instance to manage the process of deploying the target model for multiple target entities.

[0217] It should be understood that the MLEntityLoadingProcess instance has a status attribute progressStatus, and the value of this attribute can be any one of the following: running (RUNNING), cancelling (CANCELLING), suspended (SUSPENDED), finished (FINISHED), or cancelled (CANCELLED). Among them, when the value of progressStatus of the MLentityLoadingProcess instance is running (RUNNING), it can be understood that the model deployment process is in progress. Therefore, RUNNING can also be understood as "deploying" or "loading". This application does not make any limitations in this regard.

[0218] Exemplarily, when the second device creates the MLEntityLoadingProcess instance, or after creating the MLEntityLoadingProcess instance, and the value of progressStatus of this MLEntityLoadingProcess instance is RUNNING, the second device can send the above-mentioned first information to the target entity.

[0219] Optionally, the second device can report the status of the first process to the first device.

[0220] In the case where the deployment process management method is batch management, the number of the first processes can be one, and the first process corresponds to the overall process of the deployment model of multiple target entities.

[0221] In a possible implementation manner, the second device can actively report the status of the first process to the first device. The method further includes: the second device sends the third information, and the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process. Correspondingly, the first device receives the third information.

[0222] It should be understood that the third information is used to report the status of the first process. The first identification information is associated with the message carrying the identification information of the target model and has a mapping relationship with the first process. Therefore, the second device can use the first identification information or the identification information of the first process to identify the first process. The status of the first process may include: running, canceling, canceled, paused, or ended.

[0223] In another possible implementation, the second device may, in response to the second request of the first device, report the status of the first process to the first device. Specifically, before the second device sends the third information, the second device receives the second request from the first device, and this second request is used to request the status of the first process. The first request includes the first identification information, or the identification information of the first process.

[0224] Optionally, the first device may control the first process. In one possible implementation, the method further includes: the first device sends the fourth information, and the fourth information is used to control the overall process of deploying the target model by multiple target entities. The fourth information includes the first identification information and the control instruction for the first process, or the identification information of the first process and the control instruction. The control instruction includes cancel, pause, or start. Correspondingly, the second device receives this fourth information.

[0225] In the embodiments of the present application, through the model deployment method of batch management, the number of the first processes can be one. The second device can achieve the purpose of deploying the target model to multiple target entities by only maintaining one MLEntityLoadingProcess instance, which is convenient for the second device to manage the model deployment process. At the same time, in the process of the first device and the second device requesting and reporting the status of the first process and the control interaction of the first process, the process of deploying the target model by multiple target entities is regarded as a whole, without the need to pay attention to the model deployment process of each target entity, nor the need to interact with the information of each target entity, which is beneficial to reducing the signaling overhead between devices and improving the deployment efficiency of the target model.

[0226] The second method: independent management

[0227] "Independent management" can be understood as creating an independent process for management for the model deployment process of each entity among multiple target entities. In this case, the first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment process of a target entity. The identification information of the first process includes the multiple identification information corresponding to the multiple processes. Exemplarily, the second device creates multiple MLEntityLoadingProcess instances to manage the model deployment process, and these multiple MLEntityLoadingProcess instances respectively correspond to the model deployment processes of the target entities.

[0228] When the deployment process management method is independent management, when the second device reports the status of the first process to the first device, it can identify the target model deployment process corresponding to this entity through the first identification information and the identification information of a certain entity in the target entity, or it can identify the target model deployment process of a target entity through the identification information of a certain process among multiple processes.

[0229] In a possible implementation manner, the second device actively sends the fifth information, and the fifth information is used to report the status of all or part of the processes among multiple processes. The fifth information includes the first identification information, the identification information of each entity corresponding to all or part of the processes among multiple processes, and the status of each of all or part of the processes; or, the identification information of all or part of the processes, and the status of each of all or part of the processes.

[0230] In another possible implementation manner, the second device can report the status of all or part of the processes among multiple processes to the first device in response to the third request of the first device. Specifically, before the second device sends the fifth information, the second device receives the third request from the first device, and the third request is used to request the status of all or part of the processes among multiple processes. The third request includes the first identification information, the identification information of each entity corresponding to all or part of the processes among multiple processes, or the identification information of all or part of the processes.

[0231] Optionally, the first device can control multiple processes included in the first process based on a control instruction. In a possible implementation manner, the method further includes: the first device sends the sixth information, and the sixth information is used to control all or part of the processes among multiple processes. The sixth information includes the first identification information, the identification information of the entity corresponding to all or part of the processes, and the control instruction for each of all or part of the processes; or, the identification information of all or part of the processes, and the control instruction.

[0232] In the embodiments of this application, through the model deployment method of independent management, the first process can include multiple processes, and the multiple processes correspond one-to-one to the identification information of multiple target entities. The multiple processes respectively correspond to the target model deployment processes of multiple target entities. In this way, the first device can respectively obtain and control the status corresponding to the model deployment process of each target entity. In the case where some of the multiple target entities fail to be deployed, it can be determined which target entities have failed to be deployed. Based on the identification information of the target entities that have failed to be deployed, the target model can be redeployed for these target entities, which is beneficial to improving the management accuracy and enhancing the efficiency of model deployment.

[0233] In a possible implementation, the management method of the deployment process of the target model is the default one, and this default method includes batch management or independent management, that is, no explicit instruction is required from either party. Optionally, the management method of the deployment process can be agreed upon by a protocol, and the present application does not limit this.

[0234] In another possible implementation, the management method of the deployment process of the target model is indicated by the first device. The implementation method can be: the first device sends the first indication information to the second device, and this first indication information is used to indicate the management method of the deployment process of the target model.

[0235] Optionally, the first indication information can be represented by the value of the attribute infFuncBatchIndicator. In the above scenario one, the first indication information can be represented by the value of the attribute infFuncBatchIndicator of the MLEntityLoadingRequest object, but the present application does not limit this.

[0236] Optionally, the values of the first indication information infFuncBatchIndicator can include 0 and 1. When the value is 0, it indicates that the management method of the deployment process of the target model is batch management, and when the value is 1, the management method of the deployment process of the target model is independent management. The present application does not specifically limit the values of infFuncBatchIndicator and the meanings of the values.

[0237] Next, for two scenarios, namely scenario one and scenario two, two management methods for the deployment process, namely batch management and independent management, and two situations where the target entity includes all of multiple entities and the target entity includes some of multiple entities, in combination with Figures 6 to 10 a detailed description will be given.

[0238] For ease of understanding, in Figures 6 to 10 the description, taking the entity as gNB, multiple entities including the first gNB, the second gNB, and the third gNB, Figure 5 and the first device in

[0239] as NMS and the second device as EMS as examples for illustration, it should be understood that the number of entities is only for example and does not constitute a limitation to the present application.

[0240] Figure 6 This is a schematic flowchart of a model deployment method 600 provided by an embodiment of the present application. The method 600 includes:

[0241] S601. When the NMS determines that the conditions for the deployment policy of the target model are met, it sends an MLEntityLoadingRequest to the EMS. The MLEntityLoadingRequest includes the identification information of the target model and the deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingRequest.

[0242] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingRequest object shown in Table 4, and the identification information of the target model can be represented by the value of the role-related attribute mLEntityToLoadRef of the MLEntityLoadingRequest object. However, the present application does not make specific limitations on this.

[0243] Exemplarily, in this embodiment, the deployment scope of the target model indicates that the entities that need to perform inference using the target model are the first gNB, the second gNB, and the third gNB. It should be understood that the number of entities presented in this embodiment is only for illustrative purposes and does not constitute a limitation on the present application. The number of entities involved in the present application can be more or less.

[0244] S602. The EMS creates an MLEntityLoadingRequest instance and an MLEntityLoadingProcess instance.

[0245] It should be understood that an MLEntityLoadingProcess instance created by the EMS in S602 corresponds to the model deployment process of the first gNB, the second gNB, and the third gNB. That is, the process of deploying the target model for the first gNB, the second gNB, and the third gNB is regarded as a whole and managed through one process.

[0246] It should also be understood that the management method of the target model deployment process can be the default or indicated by the NMS. In the case of being indicated by the NMS, the indication information can be indicated by the infFuncBatchIndicator included in the MLEntityLoadingRequest. In this embodiment, the infFuncBatchIndicator indicates that the management method of the target model deployment process is batch management.

[0247] S603. The EMS sends a response message of the MLEntityLoadingRequest to the NMS. The response message includes the identifier of an MLEntityLoadingRequest instance and the identifier of an MLEntityLoadingProcess instance. Correspondingly, the NMS receives the response message.

[0248] S604. The NMS sends a request for the deployment process status of the target model to the EMS, including the identifier of the MLEntityLoadingRequest instance. Correspondingly, the EMS receives the request.

[0249] In another implementation of S604, the request for the deployment process status of the target model sent by the NMS to the EMS includes the identifier of the MLEntityLoadingProcess instance.

[0250] S605. The EMS reports the deployment process status of the target model to the NMS, including the identifier of the MLEntityLoadingRequest instance + the progressStatus value corresponding to the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the deployment process status of the target model.

[0251] In another implementation of S605, the EMS reports the deployment process status of the target model to the NMS, including the identifier of the MLEntityLoadingProcess instance + the progressStatus value corresponding to the MLEntityLoadingProcess instance.

[0252] S606. The NMS sends a control message for the deployment process of the target model to the EMS, including the identifier of the MLEntityLoadingRequest instance + the control instruction. Correspondingly, the EMS executes the control instruction.

[0253] In another implementation of S606, the NMS sends a control message for the deployment process of the target model to the EMS, including the identifier of the MLEntityLoadingProcess instance + the control instruction.

[0254] S607. The EMS sends the first information to the first gNB, the second gNB, and the third gNB. The first information is used to indicate the deployment of the target model. The first information includes the identification information of the target model. Correspondingly, the first gNB, the second gNB, and the third gNB receive the first information.

[0255] Optionally, the manner in which the EMS sends the first information to the first gNB, the second gNB, and the third gNB may be unicast or broadcast, and this application does not make any limitations in this regard.

[0256] It should be understood that S607 can be understood as the process in which the EMS performs model deployment and loads the target model into the first gNB, the second gNB, and the third gNB.

[0257] S608. The first gNB, the second gNB, and the third gNB deploy the target model based on the identification information of the target model.

[0258] It should be understood that when the gNB deploys the target model in S608, it can be understood as the process of the gNB activating the target model.

[0259] S609. The first gNB, the second gNB, and the third gNB perform inference based on the target model.

[0260] It should be understood that when the first gNB, the second gNB, and the third gNB need to perform inference using the target model, they each execute S609.

[0261] It should also be understood that the above S604, S605, and S606 are all optional steps.

[0262] In the embodiments of this application, by including the deployment scope of the target model in the MLEntityLoadingRequest, it is indicated that the current model deployment is for multiple gNBs. Through the model deployment method of batch management, the EMS only needs to maintain one MLEntityLoadingProcess instance to achieve the purpose of deploying the target model to multiple gNBs, which is convenient for the EMS to manage the model deployment process. At the same time, in the interaction process where the NMS requests the process status of the MLEntityLoadingProcess instance from the EMS, controls the process of the MLEntityLoadingProcess instance, or the EMS reports the process status of the MLEntityLoadingProcess instance to the NMS, only one instance needs to be targeted, which is convenient for management and also helps to reduce the signaling overhead between the interacting entities, and is conducive to improving the deployment efficiency of the target model.

[0263] The second case: Scenario 1 + The target entity includes all entities among multiple entities + Independent management

[0264] Figure 7 This is a schematic flowchart of a model deployment method 700 provided by the embodiments of this application. The method 700 includes:

[0265] S701. When NMS determines that the conditions for the deployment policy of the target model are met, it sends an MLEntityLoadingRequest to the EMS. The MLEntityLoadingRequest contains the identification information of the target model and the deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingRequest.

[0266] It should be understood that the representation methods of the identification information of the target model and the deployment scope of the target model in method 700 may be the same as those in the above method 600, and will not be elaborated here.

[0267] S702. The EMS creates an MLEntityLoadingRequest instance and multiple MLEntityLoadingProcess instances.

[0268] It should be understood that the multiple MLEntityLoadingProcess instances respectively correspond to the processes of deploying the target model for the first gNB, the second gNB, and the third gNB. Exemplarily, the number of the multiple MLEntityLoadingProcess instances is three.

[0269] It should also be understood that the management method of the deployment process of the target model can be the default or indicated by the NMS. In the case of being indicated by the NMS, the indication information can be indicated by the infFuncBatchIndicator included in the MLEntityLoadingRequest. In this embodiment, the infFuncBatchIndicator indicates that the management method of the deployment process of the target model is independent management.

[0270] S703. The EMS sends a response message of the MLEntityLoadingRequest to the NMS. The response message includes the identification of the MLEntityLoadingRequest instance and the multiple identifications corresponding to the multiple MLEntityLoadingProcess instances. Correspondingly, the NMS receives the response message.

[0271] S704. The NMS sends a status request for the MLEntityLoadingProcess instances respectively corresponding to all or part of the gNBs among the multiple gNBs to the EMS, including the identification of the MLEntityLoadingRequest instance + the identifications of all or part of the gNBs requested. Correspondingly, the EMS receives the request.

[0272] In another implementation of S704, when the NMS sends a status request for the MLEntityLoadingProcess instances corresponding to all or some of the multiple gNBs to the EMS, the status request includes the identifiers of the MLEntityLoadingProcess instances corresponding to all or some of the multiple gNBs respectively.

[0273] S705. The EMS reports the status of the MLEntityLoadingProcess instances corresponding to all or some of the multiple gNBs to the NMS, including the identifier of the MLEntityLoadingRequest instance + the identifier of all or some of the gNBs + the progressStatus value of the MLEntityLoadingProcess instances corresponding to all or some of the gNBs. Correspondingly, the NMS receives the status of the MLEntityLoadingProcess instances corresponding to all or some of the multiple gNBs.

[0274] In another implementation of S705, the EMS reports the status of the MLEntityLoadingProcess instances corresponding to all or some of the multiple gNBs to the NMS, including the identifier of the MLEntityLoadingProcess instances corresponding to all or some of the gNBs + the progressStatus value of the MLEntityLoadingProcess instances corresponding to all or some of the target gNBs.

[0275] S706. The NMS sends a control message for the target model deployment process of all or some of the multiple target gNBs to the EMS, including the identifier of the MLEntityLoadingRequest instance + the identifier of all or some of the gNBs to be controlled + the control instruction for the MLEntityLoadingProcess instance corresponding to each gNB. Correspondingly, the EMS executes the control instruction.

[0276] In another implementation of S706, the NMS sends a control message for the target model deployment process of all or some of the multiple gNBs to the EMS, including the identifier of the MLEntityLoadingProcess instances corresponding to all or some of the gNBs + the control instruction for the MLEntityLoadingProcess instance corresponding to each gNB.

[0277] S707. The EMS sends first information to the first gNB, the second gNB, and the third gNB respectively. The first information is used to indicate the deployment of the target model, and the first information includes the identification information of the target model. Correspondingly, the first gNB, the second gNB, and the third gNB receive the first information.

[0278] Optionally, the manner in which the EMS sends the first information to the first gNB, the second gNB, and the third gNB can be unicast or broadcast, and this application does not make a limitation in this regard.

[0279] S708. The first gNB, the second gNB, and the third gNB deploy the target model based on the identification information of the target model.

[0280] S709. The first gNB, the second gNB, and the third gNB perform inference based on the target model.

[0281] It should be understood that the above S704, S705, and S706 are all optional steps.

[0282] In the embodiments of this application, by independently managing the target model deployment process of multiple gNBs, the NMS can obtain and control the process status corresponding to the target model deployment of all or part of the multiple gNBs. In the case where the model deployment of some gNBs among the multiple gNBs fails, it can be determined which gNBs have failed in deployment, and the target model can be redeployed for this part of the gNBs, which is beneficial to improving the management accuracy.

[0283] The third case: Scenario 1 + The target entity includes some of the multiple entities and the number of some entities is one

[0284] Taking the target model deployed in the first gNB as an example, Figure 8 This is a schematic flowchart of a model deployment method 800 provided by the embodiments of this application. The method 800 includes:

[0285] S801. When the NMS determines that the deployment policy of the target model has been satisfied, it sends an MLEntityLoadingRequest to the EMS. The MLEntityLoadingRequest includes the identification information of the target model, the identification information of the first gNB, and the deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingRequest.

[0286] Optionally, the identification information of the first gNB can also be the identification of the inference function (AIMlInferenceFunction) corresponding to the first gNB, and this application does not make a limitation in this regard.

[0287] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingRequest object shown in Table 4. The identification information of the target model can be represented by the value of the role-related attribute mLEntityToLoadRef of the MLEntityLoadingRequest object. The identification information of the first gNB can be represented by the value of the attribute aIMlInferenceFunctionToLoadMlEntityRef, but this application does not make specific limitations in this regard.

[0288] It should be understood that in this embodiment, the MLEntityLoadingRequest includes the identification information of the first gNB. The above MLEntityLoadingRequest is used to instruct the EMS to deploy the target inference model for the first gNB. At the same time, the second gNB and the third gNB within the deployment scope of the target model are allowed to use the target model for inference.

[0289] Optionally, the first gNB is determined based on the resources and / or capabilities of multiple gNBs within the deployment scope of the target model. The NMS can query and obtain the respective resources and capabilities of multiple gNBs, and multiple gNBs can also actively report their resources and capabilities. This application does not make limitations in this regard.

[0290] In a possible implementation manner, the rule for the NMS to determine the target gNB can be: the inference resources of the first gNB are more than those of other gNBs, the inference speed of the first gNB is faster than that of other gNBs, the inference energy consumption of the first gNB is lower than that of other gNBs, the average inference result transmission delay of the first gNB within a preset time period is shorter than that of other gNBs, etc. This application does not make limitations in this regard.

[0291] S802. The EMS creates an MLEntityLoadingRequest instance and an MLEntityLoadingProcess instance.

[0292] It should be understood that in this embodiment, the MLEntityLoadingProcess instance created by the EMS is a model deployment process instance for the first gNB indicated by the NMS in S601.

[0293] S803. The EMS sends a response message of the MLEntityLoadingRequest to the NMS. The response message includes the identification of the MLEntityLoadingRequest instance and the identification of the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the response message.

[0294] S804. The NMS sends a request for the deployment process status of the target model to the EMS, including the identifier of the MLEntityLoadingRequest instance. Correspondingly, the EMS receives this request.

[0295] In another implementation of S804, the request for the deployment process status of the target model sent by the NMS to the EMS includes the identifier of the MLEntityLoadingProcess instance.

[0296] S805. The EMS reports the deployment process status of the target model to the NMS, including the identifier of the MLEntityLoadingRequest instance + the progressStatus value corresponding to the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the deployment process status of the target model.

[0297] In another implementation of S805, the EMS reports the deployment process status of the target model to the NMS, including the identifier of the MLEntityLoadingProcess instance + the progressStatus value corresponding to the MLEntityLoadingProcess instance.

[0298] S806. The NMS sends a control message for the deployment process of the target model to the EMS, including the identifier of the MLEntityLoadingRequest instance + the control instruction. Correspondingly, the EMS executes this control instruction.

[0299] In another implementation of S806, the NMS sends a control message for the deployment process of the target model to the EMS, including the identifier of the MLEntityLoadingProcess instance + the control instruction.

[0300] S807. The EMS sends the first information to the first gNB, and this first information is used to indicate the deployment of the target model. The first information includes the identification information of the target model and the deployment scope of the target model. Correspondingly, the first gNB receives this first information.

[0301] It should be understood that the identification information of the target model is used to indicate the deployment of this target model, and the deployment scope of the target model is used to indicate that any gNB belonging to this deployment scope can utilize this target model deployed on this first gNB.

[0302] S808. The first gNB deploys the target model based on the identification information of the target model.

[0303] It should be understood that the first gNB deployment target model in S808 can be understood as the process of activating the target model of the first gNB. The EMS can track the progress of activating the target model of the first gNB through the MLEntityLoadingProcess instance created in S802 above, and execute S809 while or after the first gNB activates the target model.

[0304] S809. The EMS sends second information to the second gNB and the third gNB. The second information is used to indicate that the target model has been successfully deployed on the first gNB. The second information includes the identification information of the target model and the identification information of the first gNB. Correspondingly, the second gNB and the third gNB receive the second information.

[0305] Optionally, the way for the first gNB to perform inference using the target model can be: the first gNB performs inference using the target model already deployed on itself.

[0306] Optionally, the way for the second gNB and / or the third gNB to perform inference using the target model can be: the second gNB performs inference using the target model already deployed on the first gNB, and / or, the third gNB performs inference using the target model already deployed on the first gNB, but the present application does not limit this. Exemplarily, the process for the third gNB to perform inference using the target model already deployed on the first gNB can be as in S810 - S812.

[0307] S810. The third gNB sends a first request to the first gNB. The first request is used to request the inference result of the target model. The first request includes the identification information of the target model and the input parameters of the target model. Correspondingly, the first gNB receives the first request.

[0308] It should be understood that the input parameters of the target model included in the first request sent by the third gNB to the first gNB can be the parameters of the cell served by the third gNB, and the present application does not limit this.

[0309] S811. The first gNB performs inference based on the target model and the input parameters to obtain an inference result.

[0310] It should be understood that when the first gNB determines that the third gNB belongs to the deployment scope of the target model, it inputs the input parameters of the target model included in the first request into the target model that has been deployed and completed on itself for inference to obtain an inference result.

[0311] S812. The first gNB sends the inference result to the third gNB.

[0312] It should be understood that the process for the second gNB to perform inference using the target model already deployed on the first gNB is similar to S810 - S812 above and will not be elaborated.

[0313] It should also be understood that the above S804, S805, and S806 may be optional steps.

[0314] In the embodiment of the present application, by deploying the target model to the first gNB within the deployment range of the target model, all gNBs within the deployment range can utilize the target model for inference, which is beneficial to improving the flexibility of model deployment and model usage, so as to reduce unnecessary resource waste.

[0315] The fourth case: Scenario 1 + The target entity includes some of the multiple entities and the number of some entities is not unique

[0316] In this case, the deployment process of the target model can still have two types: batch management and independent management, which can specifically include the following two:

[0317] 1) In the case of batch management, its deployment process can refer to the above method 600. The difference is that in S601, the MLEntityLoadingRequest sent by the NMS to the EMS also includes the identifiers of the first gNB, the second gNB, and the third gNB to indicate the deployment of the target model on the first gNB, the second gNB, and the third gNB. The deployment range of the target model may also include the fourth gNB, and the target model is not deployed on the fourth gNB. The process of the fourth gNB using the target model for inference can refer to S810 - S812 in method 800. The fourth gNB can request the inference result from any one of the first gNB, the second gNB, and the third gNB that has already deployed the target model. The present application does not make any limitations in this regard.

[0318] 2) In the case of independent management, its deployment process can refer to the above method 700. The difference is that in S701, the MLEntityLoadingRequest sent by the NMS to the EMS also includes the identifiers of the first gNB, the second gNB, and the third gNB to indicate the deployment of the target model on the first gNB, the second gNB, and the third gNB. The deployment range of the target model may also include the fourth gNB, and the target model is not deployed on the fourth gNB. The process of the fourth gNB using the target model for inference can refer to S810 - S812 in method 800. The fourth gNB can request the inference result from any one of the first gNB, the second gNB, and the third gNB that has already deployed the target model. The present application does not make any limitations in this regard.

[0319] The fifth case: Scenario 2 + The target entity includes all of the multiple entities + Independent management

[0320] Figure 9 Schematic flowchart of a model deployment method 900 provided by an embodiment of this application. The method 900 includes:

[0321] S901. The NMS sends an MLEntityLoadingPolicy to the EMS. The MLEntityLoadingPolicy includes the identification information of the target model, the deployment policy of the target model, and the deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingPolicy.

[0322] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingPolicy object, the deployment policy of the target model can be represented by the value of the attribute policyForLoading, and the identification information of the target model can be represented by the value of the attribute inferenceType or the value of the attribute mLEntityId. However, this application does not limit this.

[0323] S902. The EMS creates an MLEntityLoadingPolicy instance.

[0324] S903. The EMS creates multiple MLEntityLoadingProcess instances.

[0325] In a possible implementation, when the EMS determines that the first gNB, the second gNB, and the third gNB all meet the condition that the deployment policy of the target model is satisfied, it creates a first MLEntityLoadingProcess instance for deploying the target model to the first gNB, a second MLEntityLoadingProcess instance for deploying the target model to the second gNB, and a third MLEntityLoadingProcess instance for deploying the target model to the third gNB.

[0326] Optionally, under the condition that the first gNB, the second gNB, and the third gNB simultaneously meet the deployment policy of the target model, the above-mentioned first MLEntityLoadingProcess instance, second MLEntityLoadingProcess instance, and third MLEntityLoadingProcess instance can be created simultaneously; in the case where the first gNB, the second gNB, and the third gNB meet the deployment policy of the target model at different times, the above-mentioned first MLEntityLoadingProcess instance, second MLEntityLoadingProcess instance, and third MLEntityLoadingProcess instance can be created according to the order in which the first gNB, the second gNB, and the third gNB meet the deployment policy of the target model. This application does not make any limitations on this.

[0327] Exemplarily, if the first MLEntityLoadingProcess instance is created before the second MLEntityLoadingProcess instance, the time when the EMS sends the first information to the first gNB can also be earlier than the time when the EMS sends the first information to the second gNB. This application does not make any limitations on this.

[0328] In a possible implementation manner, the management method of the deployment process of the target model can be the default. In this embodiment, the management method of the deployment process of the target model is default to independent management.

[0329] In another possible implementation manner, the management method of the deployment process of the target model is indicated by the NMS, and this indication information can be indicated by the infFuncBatchIndicator included in the MLEntityLoadingPolicy. In this embodiment, the infFuncBatchIndicator indicates that the management method of the deployment process of the target model is independent management.

[0330] In yet another possible implementation manner, the management method of the deployment process of the target model is determined by the EMS. In this embodiment, the EMS determines that the management method of the deployment process of the target model is independent management.

[0331] S904. The EMS sends the identifier of the MLEntityLoadingPolicy instance and the identifiers corresponding to multiple MLEntityLoadingProcess instances to the NMS. Correspondingly, the NMS receives this message.

[0332] Optionally, the EMS may send the identifier of the MLEntityLoadingPolicy instance to the NMS immediately after executing S902, or may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S903. The identifier of the MLEntityLoadingPolicy instance and the multiple identifiers corresponding to the multiple MLEntityLoadingProcess instances may be carried by the same message or by different messages. This application does not make any restrictions on this.

[0333] S905. The NMS sends a status request for the MLEntityLoadingProcess instances corresponding to all or part of the multiple gNBs to the EMS, including the identifier of the MLEntityLoadingPolicy instance + the identifiers of all or part of the gNBs requested. Correspondingly, the EMS receives this request.

[0334] In another implementation manner of S905, in the status request sent by the NMS to the EMS for the MLEntityLoadingProcess instances corresponding to all or part of the multiple gNBs, it includes the identifiers of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs requested among the multiple gNBs.

[0335] S906. The EMS reports the status of the MLEntityLoadingProcess instances corresponding to all or part of the multiple gNBs to the NMS, including the identifier of the MLEntityLoadingPolicy instance + the identifiers of all or part of the gNBs + the progressStatus values of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs. Correspondingly, the NMS receives this status.

[0336] In another implementation manner of S906, when the EMS reports the status of the MLEntityLoadingProcess instances corresponding to all or part of the multiple gNBs to the NMS, it includes the identifiers of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs + the progressStatus values of the MLEntityLoadingProcess instances corresponding to all or part of the target gNBs.

[0337] S907. The NMS sends a control message for the target model deployment process of all or part of the gNBs in multiple target gNBs to the EMS, including the identifier of the MLEntityLoadingPolicy instance + the identifier of all or part of the gNBs to be controlled + the control instruction for the MLEntityLoadingProcess instance corresponding to each gNB. Correspondingly, the EMS executes this control instruction.

[0338] In another implementation manner of S907, the NMS sends a control message for the target model deployment process of all or part of the gNBs in multiple gNBs to the EMS, including the identifier of the MLEntityLoadingProcess instance corresponding to all or part of the gNBs + the control instruction for the MLEntityLoadingProcess instance corresponding to each gNB.

[0339] S908. The EMS sends first information to the first gNB, the second gNB, and the third gNB respectively, and this first information is used to indicate the deployment of the target model. The first information includes the identification information of the target model. Correspondingly, the first gNB, the second gNB, and the third gNB receive this first information.

[0340] S909. The first gNB, the second gNB, and the third gNB deploy the target model based on the identification information of the target model.

[0341] S910. The first gNB, the second gNB, and the third gNB perform inference based on the target model.

[0342] It should be understood that S905, S906, and S907 are optional steps.

[0343] The beneficial effects of this embodiment are similar to those of the above method 700 and will not be elaborated here.

[0344] Optionally, in Scenario 2, when the management method of the target model deployment process is determined by the EMS and the EMS monitors that the timing for each gNB to meet the target model deployment policy is the same, the batch management method can also be used to deploy the target model for multiple gNBs. In this case, in the above S903, the EMS can create only one MLEntityLoadingProcess instance, and in S905, S906, and S907, the interaction can also be completed only through the identifier of the MLEntityLoadingPolicy instance or the identifier of this MLEntityLoadingProcess instance. After the model deployment is successful, the EMS can continue to execute S908 to trigger the subsequent process, which will not be elaborated in this application.

[0345] The sixth case: Scenario 2 + The target entity includes some entities among multiple entities

[0346] Figure 10 This is a schematic flowchart of a model deployment method 1000 provided by an embodiment of the present application. The method 1000 includes:

[0347] S1001. The NMS sends the MLEntityLoadingPolicy to the EMS. The MLEntityLoadingPolicy includes the identification information of the target model, the deployment policy of the target model, and the deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingPolicy.

[0348] It should be understood that in the case where there are gNBs in the deployment scope of the target model that do not require actual deployment of the target model, the gNBs that actually need to deploy the target model can be specified by the NMS or determined by the EMS itself.

[0349] In the case where the gNBs that actually need to deploy the target model are specified by the NMS, the MLEntityLoadingPolicy may further include the identification of the gNBs that actually need to deploy the target model and are specified by the NMS. The number of gNBs that actually need to deploy the target model can be one or multiple, and the present application does not make any limitations in this regard.

[0350] In the case where the number of gNBs that actually need to deploy the target model is multiple, S1003 to S1007 in this embodiment can be replaced by S903 to S907 above. In this embodiment, taking the number of gNBs that actually need to deploy the target model as one as an example, exemplarily, the NMS specifies to deploy the target model on the first gNB. Optionally, in S1001 above, the MLEntityLoadingPolicy sent by the NMS to the EMS further includes the identification information of the first gNB.

[0351] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingPolicy object. The deployment policy of the target model can be represented by the value of the attribute policyForLoading. The identification information of the target model can also be represented by the value of the attribute inferenceType or the value of the attribute mLEntityId. The identification information of the first gNB that actually needs to deploy the target model and is specified by the NMS can be represented by the value of the attribute aIMlInferenceFunctionToLoadMlEntityRef, but the present application does not make any limitations in this regard.

[0352] S1002. The EMS creates an MLEntityLoadingPolicy instance.

[0353] S1003. The EMS creates an MLEntityLoadingProcess instance.

[0354] In a possible implementation, when the EMS determines that the deployment policy of the target model for the first gNB is satisfied, it creates an MLEntityLoadingProcess instance for deploying the target model to the first gNB.

[0355] S1004. The EMS sends the identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance to the NMS. Correspondingly, the NMS receives this message.

[0356] Optionally, the EMS may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S902, or may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S1003. The identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance may be carried by the same message or by different messages. This application does not make any restrictions on this.

[0357] S1005. The NMS sends a deployment status request for the target model to the EMS, including the identifier of the MLEntityLoadingPolicy instance. Correspondingly, the EMS receives this request.

[0358] In another implementation of S1005, when the NMS sends a deployment status request for the target model to the EMS, it includes the identifier of the MLEntityLoadingProcess instance.

[0359] S1006. The EMS reports the deployment status of the target model in multiple gNBs to the NMS, including the identifier of the MLEntityLoadingPolicy instance + the value of the progressStatus of the MLEntityLoadingProcess instance. Correspondingly, the NMS receives this status.

[0360] In another implementation of S1006, the EMS reports the deployment status of the target model to the NMS, including the identifier of the MLEntityLoadingProcess instance + the value of the progressStatus of the MLEntityLoadingProcess instance.

[0361] S1007. The NMS sends a control message of the target model deployment process to the EMS, including the identifier of the MLEntityLoadingPolicy instance + the control instruction of the MLEntityLoadingProcess instance. Correspondingly, the EMS executes this control instruction.

[0362] In another implementation of S1007, the NMS sends a control message of the target model deployment process to the EMS, including the identifier of the MLEntityLoadingProcess instance + the control instruction of the MLEntityLoadingProcess instance.

[0363] S1008. The EMS sends first information to the first gNB respectively. This first information is used to indicate the deployment of the target model, and the first information includes the identification information of the target model. Correspondingly, the first gNB receives this first information.

[0364] S1009. The first gNB deploys the target model based on the identification information of the target model.

[0365] It should be understood that the deployment of the target model by the first gNB in S1008 can be understood as the process of the first gNB activating the target model. The EMS can track the progress of the first gNB activating the target model through the MLEntityLoadingProcess instance created in the above S1002, and execute S1010 while or after the first gNB activates the target model.

[0366] S1010. The EMS sends second information to the second gNB and the third gNB. The second information is used to indicate that the target model has been successfully deployed on the first gNB. The second information includes the identification information of the target model and the identification information of the first gNB. Correspondingly, the second gNB and the third gNB receive the second information.

[0367] Optionally, the way for the first gNB to perform inference using the target model can be: the first gNB performs inference using the target model it has deployed.

[0368] Optionally, the way for the second gNB and / or the third gNB to perform inference using the target model can be: the second gNB performs inference using the target model deployed on the first gNB, and / or, the third gNB performs inference using the target model deployed on the first gNB, but the present application does not make any limitations in this regard. Exemplarily, the process for the third gNB to perform inference using the target model deployed on the first gNB can be as in S810 - S812.

[0369] S1011. The third gNB sends a first request to the first gNB. The first request is used to request the inference result of the target model, and the identification information of the target model and the input parameters of the target model are included in the first request. Correspondingly, the first gNB receives the first request.

[0370] It should be understood that the input parameters of the target model included in the first request sent by the third gNB to the first gNB may be the parameters of the cell served by the third gNB, and the present application does not limit this.

[0371] S1012. The first gNB performs inference based on the target model and the input parameters to obtain an inference result.

[0372] It should be understood that when the first gNB determines that the third gNB belongs to the deployment scope of the target model, the input parameters of the target model included in the first request are input into the target model that has been deployed and completed by itself for inference to obtain an inference result.

[0373] S1013. The first gNB sends the inference result to the third gNB.

[0374] It should be understood that the process of the second gNB using the target model deployed on the first gNB for inference is similar to the above S1010 - S1012 and will not be elaborated here.

[0375] It should be understood that S1005, S1006, and S1007 are optional steps.

[0376] The beneficial effects of this embodiment are similar to those of the above method 800 and will not be elaborated here.

[0377] It should be understood that the magnitudes of the sequence numbers of the various steps in the above embodiments do not indicate the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0378] In the above, in combination with Figures 5 to 10 , the model deployment method of the embodiments of the present application has been described in detail. Next, in combination with Figure 11 and Figure 12 , the communication device of the embodiments of the present application will be described in detail.

[0379] Figure 11 FIG. shows the communication device 1100 provided by the embodiments of the present application. The communication device 1100 includes a transceiver module 1101 and a processing module 1102.

[0380] In a first possible implementation manner, the communication device 1100 is used to implement the steps and processes corresponding to the above second device (such as EMS).

[0381] Among them, the transceiver module 1101 is used to: receive the identification information of the target model and the deployment scope of the target model, where the deployment scope of the target model is used to indicate multiple entities that need to perform inference using the target model; the processing module 1102 is used to: deploy the target model based on the identification information of the target model and the deployment scope of the target model.

[0382] Optionally, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple inference functions; or, identification information of multiple entities.

[0383] Optionally, the transceiver module 1101 is further used to: send first information to the target entity, where the first information is used to instruct the target entity to deploy the target model, and the target entity includes all or part of the multiple entities.

[0384] Optionally, the target entity includes all of the multiple entities, and the first information includes the identification information of the target model.

[0385] Optionally, the target entity includes part of the multiple entities, and the first information includes the identification information of the target model and the deployment scope of the target model.

[0386] Optionally, the transceiver module 1101 is further used to: receive the identification information of the target entity.

[0387] Optionally, the transceiver module 1101 is further used to: send second information, where the second information is used to indicate that the deployment of the target model is successful, and the second information includes the identification information of the target model and the identification information of the target entity.

[0388] Optionally, the transceiver module 1101 is further used to: send the first identification information and the identification information of the first process, where the first identification information is associated with the message carrying the identification information of the target model, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.

[0389] Optionally, the transceiver module 1101 is further used to: receive first indication information, where the first indication information is used to indicate the management method of the deployment process of the target model.

[0390] Optionally, the management method of the deployment process is batch management.

[0391] Optionally, the default management method of the deployment process of the target model is batch management.

[0392] Optionally, the first process corresponds to the overall deployment process of the target model.

[0393] Optionally, the transceiver module 1101 is further configured to: send third information, where the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.

[0394] Optionally, the transceiver module 1101 is further configured to: receive fourth information, where the fourth information is used to control the overall deployment process of the target model, and the fourth information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause, or start.

[0395] Optionally, the default management method for the deployment process of the target model is independent management.

[0396] Optionally, the deployment process management method is independent management.

[0397] Optionally, the first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of an entity.

[0398] Optionally, the transceiver module 1101 is further configured to: send fifth information, where the fifth information includes the first identification information, the identification information of each entity corresponding to all or part of the multiple processes, and the status of all or part of the processes respectively; or the identification information of all or part of the processes, and the status of all or part of the processes respectively, and the status includes: running, canceling, canceled, paused, or ended.

[0399] Optionally, the transceiver module 1101 is further configured to: receive sixth information, where the sixth information is used to control all or part of the multiple processes, and the sixth information includes the first identification information, the identification information of the entity corresponding to all or part of the processes, and a control instruction for each of all or part of the processes; or the identification information of all or part of the processes, and the control instruction.

[0400] Optionally, the identification information of the target model is sent through a request message or a configuration message.

[0401] Optionally, the deployment scope of the target model is sent through a request message or a configuration message.

[0402] In the second possible implementation manner, the communication device 1100 is configured to implement the steps and processes corresponding to the above first device (such as NMS).

[0403] Among them, the processing module 1102 is configured to: determine the identification information of the target model and the deployment scope of the target model, where the deployment scope of the target model is used to indicate multiple entities that need to perform inference using the target model; the transceiver module 1101 is configured to: send the identification information of the target model and the deployment scope of the target model.

[0404] Optionally, the deployment scope of the target model includes one or more of the following: information on geographical regions; identification information on network regions; identification information on multiple inference functions; or identification information on multiple entities.

[0405] Optionally, the transceiver module 1101 is further configured to: send the identification information of a target entity, where the target entity includes some entities among the multiple entities, and the target entity is determined based on the resources and / or capabilities of the multiple entities.

[0406] Optionally, the transceiver module 1101 is further configured to: send the first identification information and the identification information of a first process, where the first identification information is associated with a message carrying the identification information of the target model, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.

[0407] Optionally, the transceiver module 1101 is further configured to: receive first indication information, where the first indication information is used to indicate the management method of the deployment process of the target model.

[0408] Optionally, the management method of the deployment process is batch management.

[0409] Optionally, the default management method of the deployment process of the target model is batch management.

[0410] Optionally, the first process corresponds to the overall deployment process of the target model.

[0411] Optionally, the transceiver module 1101 is further configured to: send third information, where the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.

[0412] Optionally, the transceiver module 1101 is further configured to: receive fourth information, where the fourth information is used to control the overall deployment process of the target model, and the fourth information includes the first identification information and a control instruction for the first process, or the identification information of the first process and a control instruction, and the control instruction includes cancel, pause, or start.

[0413] Optionally, the default management method of the deployment process of the target model is independent management.

[0414] Optionally, the management method of the deployment process is independent management.

[0415] Optionally, the first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of one entity.

[0416] Optionally, the transceiver module 1101 is further configured to: send a fifth piece of information, where the fifth piece of information includes first identification information, identification information of each entity corresponding to all or part of multiple processes, and the status of all or part of the processes respectively; or, identification information of all or part of the processes, and the status of all or part of the processes respectively, where the status includes: running, canceling, canceled, paused, or ended.

[0417] Optionally, the transceiver module 1101 is further configured to: receive a sixth piece of information, where the sixth piece of information is used to control all or part of multiple processes, and the sixth piece of information includes first identification information, identification information of the entity corresponding to all or part of the processes, and a control instruction for each of all or part of the processes; or, identification information of the entity corresponding to all or part of the processes, and a control instruction.

[0418] Optionally, the identification information of the target model is sent through a request message or a configuration message.

[0419] Optionally, the deployment scope of the target model is sent through a request message or a configuration message.

[0420] In a third possible implementation manner, the communication device 1100 is configured to implement the steps and processes corresponding to the above third device (such as an entity, gNB).

[0421] Among them, the transceiver module 1101 is configured to: receive first information, where the first information is used to indicate the deployment of the target model, and the first information includes the deployment scope of the target model and the identification information of the target model, and the deployment scope of the target model is used to indicate multiple entities that need to perform inference using the target model; the processing module 1102 is configured to: deploy the target model based on the first information.

[0422] Optionally, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple inference functions; or, identification information of multiple entities.

[0423] Optionally, the transceiver module 1101 is further configured to: receive a first request, where the first request is used to request the inference result of the target model, and the first request includes the identification information of the target model and the input parameters of the target model; the processing module 1102 is further configured to: obtain the inference result through the target model based on the input parameters.

[0424] Optionally, the transceiver module 1101 is further configured to: send the inference result.

[0425] Optionally, the transceiver module 1101 is further configured to: receive second information, where the second information is used to indicate that the target model is successfully deployed, and the second information includes the identification information of the target model and the identification information of the target entity.

[0426] It should be understood that the device 1100 here is embodied in the form of functional modules. The term "module" here may refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a proprietary processor or a group of processors, etc.) for executing one or more software or firmware programs, and a memory, a combined logic circuit and / or other suitable components that support the described functions. In an alternative example, those skilled in the art can understand that the device 1100 can specifically be the first device, the second device or the third device in the above embodiments, or, the functions described in the above embodiments can be integrated in the device 1100, and the device 1100 can be used to execute each process and / or step corresponding to the first device, the second device or the third device in the above method embodiments. To avoid repetition, it will not be elaborated here.

[0427] The above device 1100 has the function of implementing the corresponding steps executed by the first device, the second device or the third device in the above method; the above function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0428] In the embodiments of the present application, Figure 11 the device 1100 in

[0429] Figure 12 Fig. shows a schematic block diagram of a communication device 1200 provided by an embodiment of the present application. The device 1200 includes a processor 1201, a transceiver 1202 and a memory 1203. Among them, the processor 1201, the transceiver 1202 and the memory 1203 communicate with each other through an internal connection path. The memory 1203 is used to store instructions, and the processor 1201 is used to execute the instructions stored in the memory 1203 to control the transceiver 1202 to send signals and / or receive signals.

[0430] It should be understood that the device 1200 may specifically be the first device, the second device, or the third device in the foregoing embodiments, and may be used to execute the respective steps and / or processes corresponding to the first device, the second device, or the third device in the foregoing method embodiments. Optionally, the memory 1203 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may further include a non-volatile random access memory. For example, the memory may further store information about the device type. The processor 1201 may be used to execute the instructions stored in the memory, and when the processor 1201 executes the instructions stored in the memory, the processor 1201 is used to execute the respective steps and / or processes of the foregoing method embodiments. The transceiver 1202 may include a transmitter and a receiver. The transmitter may be used to implement the respective steps and / or processes corresponding to the foregoing transceiver for performing a sending action, and the receiver may be used to implement the respective steps and / or processes corresponding to the foregoing transceiver for performing a receiving action.

[0431] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0432] In the implementation process, the respective steps of the foregoing method may be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware processor, or executed and completed by a combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor executes the instructions in the memory and combines its hardware to complete the steps of the foregoing method. To avoid repetition, it will not be described in detail here.

[0433] The present application further provides a computer-readable storage medium, which is used to store a computer program for implementing the method shown in the foregoing method embodiments.

[0434] The present application further provides a computer program product, which includes computer program code (which may also be referred to as a computer program, an instruction). When the computer program runs on a computer, the computer may execute the method shown in the foregoing method embodiments.

[0435] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0436] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0437] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0438] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0439] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0440] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.< / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc>

Claims

1. A model deployment method, characterized in that, The method includes: Receiving the identification information of the target model and the deployment scope of the target model, where the deployment scope of the target model is used to indicate multiple entities that need to perform inference using the target model; Deploying the target model based on the identification information of the target model and the deployment scope of the target model.

2. The method according to claim 1, wherein The deployment scope of the target model includes one or more of the following: Information of a geographical area; Identification information of a network area; Identification information of multiple inference functions; or, Identification information of multiple entities.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Sending the first identification information and the identification information of the first process, where the first identification information is associated with a message carrying the identification information of the target model, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.

4. The method according to claim 3, characterized in that The management method of the deployment process of the target model is batch management.

5. The method according to claim 4, wherein The first process corresponds to the overall deployment process of the target model.

6. The method according to claim 4 or 5, characterized in that, The method further includes: Sending the third information, where the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.

7. The method according to any one of claims 4 to 6, characterized in that The method further includes: Receiving the fourth information, where the fourth information is used to control the overall deployment process of the target model, and the fourth information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause, or start.

8. The method according to claim 3, wherein The management method of the deployment process of the target model is independent management.

9. The method according to claim 8, wherein The first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of one entity.

10. The method according to claim 8 or 9, characterized in that The method further includes: Sending the fifth information, where the fifth information includes the first identification information, the identification information of each entity corresponding to all or part of the multiple processes, and the status of each of all or part of the processes; or, The identification information of all or part of the processes, and the status of each of all or part of the processes, and the status includes: running, canceling, canceled, paused, or ended.

11. The method according to any one of claims 8 to 10, characterized in that, The method further includes: Receiving the sixth information, where the sixth information is used to control all or part of the multiple processes, and the sixth information includes the first identification information, the identification information of the entity corresponding to all or part of the processes, and a control instruction for each of all or part of the processes; or, The identification information of all or part of the processes, and the control instruction.

12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: Receiving the first indication information, where the first indication information is used to indicate the management method of the deployment process of the target model.

13. The method according to any one of claims 4 to 7, characterized in that The default management method of the deployment process of the target model is batch management.

14. The method according to any one of claims 8 to 11, characterized in that The default management method of the deployment process of the target model is independent management.

15. A model deployment method, characterized in that, The method includes: Determining the identification information of the target model and the deployment scope of the target model, where the deployment scope of the target model is used to indicate multiple entities that need to perform inference using the target model; Send the identification information of the target model and the deployment scope of the target model.

16. The method according to claim 15, wherein The deployment scope of the target model includes one or more of the following: Information of a geographical area; Identification information of a network area; Identification information of multiple inference functions; or, Identification information of multiple entities.

17. The method according to claim 15 or 16, characterized in that The method further includes: Send the first identification information and the identification information of the first process, where the first identification information is associated with the message carrying the identification information of the target model, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.

18. The method according to claim 17, wherein The deployment process of the target model is managed in a batch manner.

19. The method according to claim 18, wherein The first process corresponds to the overall deployment process of the target model.

20. The method according to claim 18 or 19, characterized in that, The method further includes: Send the third information, where the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, and the status includes: running, canceling, canceled, paused, or ended.

21. The method according to any one of claims 18 to 20, characterized in that, The method further includes: Receive the fourth information, where the fourth information is used to control the overall deployment process of the target model, and the fourth information includes the first identification information and the control instruction for the first process, or the identification information of the first process and the control instruction, and the control instruction includes cancel, pause, or start.

22. The method according to claim 17, wherein The deployment process of the target model is managed independently.

23. The method according to claim 22, wherein The first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of one entity.

24. The method according to claim 22 or 23, characterized in that, The method further includes: Send the fifth information, where the fifth information includes the first identification information, the identification information of each entity corresponding to all or part of the multiple processes, and the status of each of the all or part of the processes; or, The identification information of the all or part of the processes, and the status of each of the all or part of the processes, and the status includes: running, canceling, canceled, paused, or ended.

25. The method according to any one of claims 22 to 24, characterized in that, The method further includes: Receive the sixth information, where the sixth information is used to control all or part of the multiple processes, and the sixth information includes the first identification information, the identification information of the entity corresponding to the all or part of the processes, and the control instruction for each of the all or part of the processes; or, The identification information of the entity corresponding to the all or part of the processes, and the control instruction.

26. The method according to any one of claims 15 to 25, characterized in that The method further includes: Receive the first indication information, where the first indication information is used to indicate the management method of the deployment process of the target model.

27. The method according to any one of claims 18 to 21, characterized in that, The default management method of the deployment process of the target model is batch management.

28. The method according to any one of claims 22 to 25, characterized in that The default management method of the deployment process of the target model is independent management.

29. The method according to any one of claims 1 to 28, characterized in that, The deployment scope of the target model is sent through a request message or a configuration message.

30. A communication device, characterized in that, Includes: A module for executing the method according to any one of claims 1 to 14, or a module for executing the method according to any one of claims 15 to 29.

31. A communication device, characterized in that, Includes: A processor, the processor being coupled to a memory that stores computer-executable instructions, the processor executing the computer-executable instructions stored in the memory such that the processor performs the method according to any one of claims 1 to 14, or performs the method according to any one of claims 15 to 29.

32. A communication system, characterized in that, Comprising a first device and a second device, wherein the second device is configured to perform the method according to any one of claims 1 to 14, and the first device is configured to perform the method according to any one of claims 15 to 29.

33. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program comprising instructions for implementing the method according to any one of claims 1 to 14, or for performing the method according to any one of claims 15 to 29.

34. A computer program product, comprising computer program code, characterized in that, When the computer program code runs on a computer, causing the computer to implement the method according to any one of claims 1 to 14, or to perform the method according to any one of claims 15 to 29.

Citation Information

Cited By

  • Model deployment method and communication device

    WO2025130688A1