Model deployment method and communication device
By receiving the identification information and deployment scope of the target model, the target model of multiple entities can be deployed at one time, solving the problems of high complexity and low efficiency of multiple entities in the prior art, and improving the model deployment efficiency.
Patent Information
- Application Number
- PCT/CN2024/137944
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-09
- Publication Date
- 2025-06-26
AI Technical Summary
The prior art has high complexity and low efficiency when deploying target models on multiple entities, and cannot effectively manage the model deployment process of multiple entities.
By receiving the identification information and deployment scope of the target model, the target model is deployed in one go to multiple entities based on this information, reducing deployment complexity and improving efficiency.
The target model deployment process for multiple entities is simplified, the model deployment efficiency is improved, and signaling overhead is reduced.
Smart Images

Figure CN2024137944_26062025_PF_FP_ABST
Abstract
Description
Model deployment method and communication device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 21, 2023, with application number 202311777266.2 and application name “Model Deployment Method and Communication Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communications, and in particular to a model deployment method and a communication device. Background Art
[0003] To improve network intelligence and automation, artificial intelligence (AI) and machine learning (ML) technologies are increasingly being applied. For example, access network equipment can include a service module that implements functions based on models (such as reasoning). These modules can then implement these functions based on the models, enhancing the intelligence of the access network equipment.
[0004] After the model has gone through the training phase and / or simulation phase, it also needs to go through the deployment phase to be deployed on the entity that undertakes the reasoning function in order to perform the reasoning function on the entity. The model deployment process can be understood as the process of making the model available on the entity that undertakes the reasoning function. In some examples, a model deployment process can only enable one entity (such as an access network device) to have the ability to obtain the model reasoning results, but in some scenarios, there are multiple entities with reasoning requirements. The existing method requires multiple model deployment processes for multiple entities, which is inefficient. Summary of the Invention
[0005] The present application provides a model deployment method and a communication device, which are beneficial to reducing the complexity of deploying target models for multiple entities and improving the efficiency of model deployment.
[0006] In the first aspect, the present application provides a model deployment method, which can be executed by a second device or a component in the second device (such as a chip, a chip system, etc.), and the second device has the ability to provide model loading services. The second device can be a network management system, a network element management system, a network device or a wireless access network device, or any other device or system that can provide model loading services. This application does not specifically limit this. The method includes: receiving identification information of a target model and a deployment scope of the target model, the deployment scope of the target model is used to indicate multiple entities that need to use the target model for reasoning; deploying the target model based on the identification information of the target model and the deployment scope of the target model.
[0007] In an embodiment of the present application, the information received by the second device includes the deployment scope of the target model, which is used to indicate multiple entities that need to use the target model for reasoning. Based on this deployment scope, the second device can deploy the target model for all or part of the multiple entities indicated by the deployment scope. Compared with some examples, in scenarios where models need to be deployed for multiple entities, the second device needs to receive identification information of the target model multiple times for multiple entities. The method provided in the present application indicates the deployment scope of the target model during the model deployment process, so that one model deployment process can correspond to multiple entities within the deployment scope that need to use the target model for reasoning, reducing the complexity of deploying target models for multiple entities and helping to improve model deployment efficiency.
[0008] In combination with the first aspect, in some implementations of the first aspect, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple reasoning functions; or identification information of multiple entities.
[0009] It should be understood that the deployment scope of the target model may refer to a scope that includes multiple entities that need to use the target model for reasoning, which may be a substantial regional scope or an abstract scope, and this application does not limit this.
[0010] Optionally, the information about the geographic area may be administrative division information, longitude and latitude information, etc., which is not limited in this application. It is understood that a first mapping relationship may exist between the information about the geographic area and the entity that performs the reasoning function, and the second device may maintain the first mapping relationship. In one possible implementation, the second device may, based on the information about the geographic area and the first mapping relationship, determine multiple entities in the geographic area as entities that need to be reasoned using the target model. Furthermore, the second device may perform subsequent steps for the determined multiple entities.
[0011] Optionally, the identification information of the network area can be identification information of any subnet that can distinguish the network area, such as a local area network, a metropolitan area network, or a wide area network, and this application does not limit this. It is understood that a second mapping relationship can exist between the identification information of the network area and the entity that performs the reasoning function, and the second device can maintain this second mapping relationship. In one possible implementation, the second device can determine multiple entities that need to be reasoned using the target model based on the identification information of the network area and the second mapping relationship, and further perform subsequent steps on the determined multiple entities.
[0012] In some examples, an entity that performs an inference function may be referred to as an inference function (AIM1InferenceFunction). Optionally, identification information of multiple inference functions may be used to identify multiple entities that perform inference functions. For example, the deployment scope of the target model may be indicated by multiple AIM1InferenceFunction identifiers or an AIM1InferenceFunction list.
[0013] Optionally, the identification information of the multiple entities can be any information that can uniquely identify the entities, such as the Internet Protocol (IP) addresses, Media Access Control (MAC) addresses, etc. of the multiple entities, and this application does not limit this. In some examples, the identification information of the multiple entities can be presented in the form of a list, but this application does not specifically limit this.
[0014] In combination with the first aspect, in certain implementations of the first aspect, deploying the target model based on the identification information of the target model and the deployment scope of the target model includes: sending first information to the target entity, the first information being used to instruct the target entity to deploy the target model, and the target entity includes all or part of the multiple entities.
[0015] In a possible implementation, the target entity is determined by the second device from a deployment scope of the target model, and the target entity is determined by the second device based on resources and / or capabilities of multiple entities.
[0016] In another possible implementation manner, the target entity is specified by the first device, and the target entity is determined by the first device based on resources and / or capabilities of multiple entities.
[0017] In combination with the first aspect, in some implementations of the first aspect, the target entity includes all entities in the multiple entities, and the first information includes identification information of the target model.
[0018] In this embodiment, the multiple entities indicated by the deployment scope of the target model are all target entities. Each of the multiple entities may perform reasoning using the target model by performing reasoning based on its deployed target model and obtaining a reasoning result. In this case, the first information sent by the second device to the target entity includes identification information of the target model, thereby instructing the target entity to deploy the target model.
[0019] In combination with the first aspect, in some implementations of the first aspect, the target entity includes some entities among the multiple entities, and the first information includes identification information of the target model and a deployment scope of the target model.
[0020] In this embodiment, the multiple entities indicated by the deployment scope of the target model include the target entity and other entities except the target entity. In this case, the target entity can perform reasoning based on the target model deployed by itself, and other entities can perform reasoning using the target model deployed on the target entity.
[0021] In combination with the first aspect, in some implementations of the first aspect, when the target entity is specified by the first device, the method further includes: receiving identification information of the target entity.
[0022] In combination with the first aspect, in some implementations of the first aspect, the method further includes: sending second information, where the second information is used to indicate that the target model is successfully deployed, and the second information includes identification information of the target model and identification information of the target entity.
[0023] It should be understood that after the second device successfully deploys the target model on the target entity, it can notify the first entity (an entity other than the target entity among multiple entities) of the successful deployment of the target model through the second information. The identification information of the target model in the second information is used to indicate the successfully deployed model, and the identification information of the target entity is used to indicate on which entity the target model is deployed, and the deployment scope of the target model can be used by the target entity to authenticate the first entity.
[0024] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: sending first identification information and identification information of a first process, the first identification information being associated with a message carrying the identification information of the target model, the first identification information being mapped to the first process, and the first process being used to manage the deployment process of the target model.
[0025] Optionally, the first identification information and the identification information of the first process may be sent via a response message, and the response message may be a response to a message carrying the identification information of the target model, but this application does not limit this.
[0026] In combination with the first aspect, in some implementations of the first aspect, the method further includes: receiving first indication information, where the first indication information is used to indicate a deployment process management method of the target model.
[0027] In combination with the first aspect, in some implementations of the first aspect, the deployment process management method can be batch management.
[0028] “Batch management” can be understood as treating the target model deployment process for multiple target entities as a whole. In this case, the first process corresponds to the overall deployment process of the target model.
[0029] In one possible implementation, the first indication information indicates that the deployment process management mode is batch management. In another possible implementation, the deployment process management mode of the target model defaults to batch management. When the deployment process management mode is batch management, the first process corresponds to the overall deployment process of the target model.
[0030] In combination with the first aspect, in some implementations of the first aspect, the second device may report the overall deployment progress of the target model.
[0031] In one possible implementation, the second device actively reports the status of the first process, and the method further includes: sending third information, the third information including the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, the status including: running, canceling, canceled, paused or ended.
[0032] In another possible implementation, the second device reports the status of the first process in response to the request of the first device. Before the second device sends the third information, the method also includes: the second device receives a second request from the first device, where the second request is used to request the status of the first process, and the first request includes first identification information, or identification information of the first process.
[0033] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: receiving fourth information, the fourth information being used to control the overall deployment process of the target model, the fourth information including the first identification information and control instructions for the first process, or the identification information of the first process and the control instructions, the control instructions including canceling, pausing or starting.
[0034] In the embodiment of the present application, through the batch management model deployment method, the number of the first process can be one, and the second device can maintain only one process instance to achieve the purpose of deploying the target model to multiple target entities, which facilitates the second device to manage the model deployment process. At the same time, during the state request and report of the first process and the control interaction between the first device and the second device, the process of deploying the target model to multiple target entities is regarded as a whole. There is no need to pay attention to the model deployment process of each target entity, nor is there any need to exchange information of each target entity. This is conducive to reducing the signaling overhead between the various devices and improving the deployment efficiency of the target model.
[0035] In combination with the first aspect, in certain implementations of the first aspect, the deployment process management of the target model may also be independent management.
[0036] "Independent management" can be understood as creating an independent process to manage the model deployment process of each entity in multiple target entities.
[0037] In one possible implementation, the first indication information indicates that the deployment process management mode of the target model is independent management by default; in another possible implementation, the deployment process management mode is independent management. When the deployment process management mode is independent management, the first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of an entity.
[0038] In combination with the first aspect, in certain implementations of the first aspect, the second device may report the status of all or part of the multiple processes.
[0039] In one possible implementation, the second device actively reports the status of all or part of the multiple processes, and the method further includes: sending fifth information, the fifth information including the first identification information, identification information of each entity corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, identification information of all or part of the processes, and the respective status of all or part of the processes, the status including: running, canceling, canceled, paused or ended.
[0040] In another possible implementation, the second device reports the status of all or some of the multiple processes to the first device in response to the third request from the first device. Specifically, before the second device sends the fifth information, the second device receives a third request from the first device, where the third request is used to request the status of all or some of the multiple processes, and the third request includes the first identification information, identification information of each entity corresponding to all or some of the multiple processes, or identification information of all or some of the processes.
[0041] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: receiving sixth information, the sixth information being used to control all or part of the multiple processes, the sixth information including the first identification information, identification information of the entities corresponding to the all or part of the processes, and control instructions for each of the all or part of the processes; or, identification information of the all or part of the processes, and the control instructions.
[0042] In an embodiment of the present application, through an independently managed model deployment method, the first process may include multiple processes, and the multiple processes correspond one-to-one to the identification information of multiple target entities, and the multiple processes correspond to the target model deployment processes of the multiple target entities. In this way, the first device can respectively obtain and control the status corresponding to the model deployment process of each target entity. In the case that some target entities among the multiple target entities fail to be deployed, it can be determined which target entities failed to be deployed, and based on the identification information of the target entities that failed to be deployed, the target models can be re-deployed for these target entities, which is conducive to improving management accuracy and improving the efficiency of model deployment.
[0043] On the second aspect, the present application further provides a model deployment method, which can be executed by a first device or a component in the first device (such as a chip, a chip system, etc.), the first device has the ability to call a model loading service, and the first device can be a network management system, a network element management system, a network device or a wireless access network device, or any other device or system that has the ability to call a model loading service. This application does not specifically limit this. The method includes: determining the identification information of the target model and the deployment scope of the target model, the deployment scope of the target model is used to indicate multiple entities that need to use the target model for reasoning; sending the identification information of the target model and the deployment scope of the target model.
[0044] In combination with the second aspect, in some implementations of the second aspect, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple reasoning functions; or identification information of multiple entities.
[0045] In combination with the second aspect, in some implementations of the second aspect, the method further includes: sending identification information of the target entity, the target entity including some entities among the multiple entities, and the target entity is determined based on the resources and / or capabilities of the multiple entities.
[0046] In combination with the second aspect, in certain implementations of the second aspect, the method further includes: sending first identification information and identification information of a first process, the first identification information being associated with a message carrying the identification information of the target model, the first identification information being mapped to the first process, and the first process being used to manage the deployment process of the target model.
[0047] In combination with the second aspect, in some implementations of the second aspect, the method further includes: receiving first indication information, where the first indication information is used to indicate a deployment process management method of the target model.
[0048] In combination with the second aspect, in some implementations of the second aspect, the deployment process management method is batch management.
[0049] In combination with the second aspect, in certain implementations of the second aspect, the deployment process management mode of the target model is batch management by default.
[0050] In combination with the second aspect, in some implementations of the second aspect, the first process corresponds to the overall deployment process of the target model.
[0051] In combination with the second aspect, in some implementations of the second aspect, the method also includes: sending third information, the third information including the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, the status including: running, canceling, canceled, paused or ended.
[0052] In combination with the second aspect, in certain implementations of the second aspect, the method further includes: receiving fourth information, the fourth information being used to control the overall deployment process of the target model, the fourth information including the first identification information and control instructions for the first process, or the identification information of the first process and the control instructions, the control instructions including canceling, pausing or starting.
[0053] In combination with the second aspect, in certain implementations of the second aspect, the deployment process management mode of the target model is independent management by default.
[0054] In combination with the second aspect, in some implementations of the second aspect, the deployment process management method is independent management.
[0055] In combination with the second aspect, in some implementations of the second aspect, the first process includes multiple processes, and any one of the multiple processes corresponds to a model deployment of an entity.
[0056] In combination with the second aspect, in some implementations of the second aspect, the method further includes: sending fifth information, the fifth information including the first identification information, identification information of each entity corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, identification information of all or part of the processes, and the respective status of all or part of the processes, the status including: running, canceling, canceled, paused or ended.
[0057] In combination with the second aspect, in certain implementations of the second aspect, the method further includes: receiving sixth information, the sixth information being used to control all or part of the multiple processes, the sixth information including the first identification information, identification information of the entity corresponding to all or part of the processes, and control instructions for each process in all or part of the processes; or, identification information of the entity corresponding to all or part of the processes, and the control instructions.
[0058] In combination with the second aspect, in some implementations of the second aspect, the identification information of the target model is sent via a request message or a configuration message.
[0059] In combination with the second aspect, in some implementations of the second aspect, the deployment scope of the target model is sent via a request message or a configuration message.
[0060] On the third aspect, the present application also provides a model deployment method, which can be executed by a third device or a component in the third device (such as a chip, a chip system, etc.). The third device has the ability to assume the reasoning function. The third device can be a network management system, a network element management system, a network device or a wireless access network device, or other entities, devices or systems that can assume the reasoning function. This application does not make specific restrictions on this. It includes: receiving first information, the first information is used to indicate the deployment target model, the first information includes the deployment scope of the target model and the identification information of the target model, the deployment scope of the target model is used to indicate multiple entities that need to use the target model for reasoning; based on the first information, deploying the target model.
[0061] In combination with the third aspect, in certain implementations of the third aspect, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple reasoning functions; or identification information of multiple entities.
[0062] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: receiving a first request, the first request being used to request the inference result of the target model, the first request including identification information of the target model and input parameters of the target model; and obtaining the inference result through the target model based on the input parameters.
[0063] In combination with the third aspect, in some implementations of the third aspect, the method further includes: sending the inference result.
[0064] In combination with the third aspect, in certain implementations of the third aspect, the method further includes: receiving second information, the second information being used to indicate that the target model is successfully deployed, the second information including identification information of the target model and identification information of the target entity.
[0065] In a fourth aspect, the present application provides a communication device, comprising: a module for executing the method as described in any one of the first aspect, the second aspect or the third aspect.
[0066] In a fifth aspect, the present application provides another communication device, comprising a processor coupled to a memory and configured to execute instructions in the memory to implement the method of any possible implementation of the first, second, or third aspects described above. Optionally, the device further comprises a memory. Optionally, the device further comprises a communication interface, the processor coupled to the communication interface.
[0067] In a sixth aspect, the present application provides a communication system comprising a first device, a second device, and a third device, wherein the second device is used to execute any method in the first aspect, the first device is used to execute any method in the second aspect, and the third device is used to execute any method in the third aspect.
[0068] In a seventh aspect, the present application provides a processor comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is configured to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method of any possible implementation of the first, second, or third aspects described above.
[0069] In a specific implementation, the processor may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be a transistor, a gate circuit, a trigger, or various logic circuits. The input signal received by the input circuit may be, for example, but not limited to, received and input by a receiver, and the signal output by the output circuit may be, for example, but not limited to, output to and transmitted by a transmitter. The input circuit and the output circuit may be the same circuit, which functions as an input circuit and an output circuit at different times. The embodiments of the present application do not limit the specific implementation of the processor and various circuits.
[0070] In an eighth aspect, a processing device is provided, comprising a processor and a memory. The processor is configured to read instructions stored in the memory and receive signals via a receiver and transmit signals via a transmitter to execute the method of any possible implementation of the first, second, or third aspects.
[0071] Optionally, there are one or more processors and one or more memories.
[0072] Optionally, the memory may be integrated with the processor, or the memory may be provided separately from the processor.
[0073] In the specific implementation process, the memory can be a non-transitory memory, such as a read-only memory (ROM), which can be integrated with the processor on the same chip or can be set on different chips. The embodiments of the present application do not limit the type of memory and the setting method of the memory and the processor.
[0074] It should be understood that related data interaction processes, such as sending indication information, can be the process of outputting indication information from the processor, and receiving capability information can be the process of receiving input capability information from the processor. Specifically, the output data of the processing can be output to the transmitter, and the input data received by the processor can come from the receiver. The transmitter and receiver can be collectively referred to as a transceiver.
[0075] The processing device in the above-mentioned eighth aspect can be a chip. The processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading the software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0076] In the ninth aspect, a computer program product is provided, which includes: a computer program (also referred to as code, or instructions), which, when run, enables the computer to execute the method in any possible implementation of the first, second or third aspects above.
[0077] In the tenth aspect, a computer-readable storage medium is provided, which stores a computer program (also referred to as code, or instructions) which, when run on a computer, enables the computer to execute a method in any possible implementation of the first, second or third aspects above. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] FIG1 is a model workflow provided by an embodiment of the present application;
[0079] FIG2 is a working diagram of a network resource model provided in an embodiment of the present application;
[0080] FIG3 is a working diagram of a network resource model provided by an embodiment of the application;
[0081] FIG4 is a schematic diagram of a communication system applicable to the embodiment of the application;
[0082] FIG5 is a schematic flow chart of a model deployment method provided in an embodiment of the present application;
[0083] FIG6 is a schematic flow chart of a model deployment method provided in an embodiment of the present application;
[0084] FIG7 is a schematic flow chart of another model deployment method provided in an embodiment of the present application;
[0085] FIG8 is a schematic flow chart of another model deployment method provided in an embodiment of the present application;
[0086] FIG9 is a schematic flow chart of another model deployment method provided in an embodiment of the present application;
[0087] FIG10 is a schematic flow chart of a model deployment method provided in an embodiment of the present application;
[0088] FIG11 is a schematic block diagram of a communication device provided in an embodiment of the present application;
[0089] FIG12 is a schematic block diagram of another communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0090] The technical solution in this application will be described below with reference to the accompanying drawings.
[0091] Before introducing this application, the following points are explained.
[0092] First, in the embodiments described below, various terms and abbreviations, such as baseline data and differential data, are provided for ease of description and should not be construed as limiting this application. This application does not exclude the possibility of defining other terms in existing or future protocols that can achieve the same or similar functions.
[0093] Second, the first, second and various numerical numbers in the embodiments shown below are only used for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
[0094] Third, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c, where a, b, c can be single or multiple.
[0095] In order to improve the intelligence and automation level of the network, artificial intelligence (AI) and machine learning (ML) technologies are gradually being applied to promote network intelligence.
[0096] Figure 1 exemplifies the AI / ML workflow, which mainly includes the training phase, simulation phase, deployment phase and inference phase. It should be understood that, in general, the AI / ML workflow is implemented in sequence according to the training phase, simulation phase, deployment phase and inference phase, but in some possible implementations, the AI / ML model may execute the inference phase after the training phase and / or simulation phase, or skip the simulation phase after the training phase to execute the deployment phase and inference phase, or enter the training phase after the simulation phase or inference phase. Among them, the deployment phase is the process of loading the model into the inference function. In one possible implementation, after the model has gone through the training phase and the simulation phase, it is loaded into the corresponding inference function through the deployment phase to enter the inference phase to execute the inference process.
[0097] In some examples, the role that provides the ML entity loading management service (MnS) is called a producer, and the role that calls the ML entity loading management service is called a consumer. Two sub-use cases and solutions for these two sub-use cases are defined for the deployment phase.
[0098] The two sub-use cases include:
[0099] Sub-use case 1: Consumer requested ML entity loading.
[0100] Sub-use case 2: The consumer configures the policy for the producer, and the producer triggers the model loading (control of producer-initiated ML entity loading).
[0101] Solutions for two sub-use cases:
[0102] The current solution defines three information object classes (IOCs): MLEntityLoadingRequest, MLEntityLoadingPolicy, and MLEntityLoadingProcess, which represents the ML model loading process. MLEntityLoadingRequest and MLEntityLoadingProcess are used in the scenario described in sub-use case 1, while MLEntityLoadingPolicy and MLEntityLoadingProcess are used in the scenario described in sub-use case 2.
[0103] The following describes in detail the attributes of the above three information object classes.
[0104] 1.MLEntityLoadingRequest< <ioc>>
[0105] MLEntityLoadingRequest< <ioc>> represents an ML model loading request created by a consumer. The consumer uses this IOC to request the producer to load the ML model into the target inference function. The properties of this IOC are shown in Table 1.
[0106] Table 1
[0107] The requestStatus attribute indicates the status of the request and can be any of the following: NOT_STARTED, LOADING_IN_PROGRES, SUSPENDED, FINISHED_SUCCESS, FINISHED_FAILED, or CANCELLED. The cancelRequest attribute indicates whether the consumer canceled the ML entity loading request and can be TRUE or FALSE. The mLEntityToLoadRef role attribute indicates the identifier of the model to be loaded in the ML model loading request created by the consumer.
[0108] It should be understood that the "support qualifier" can take the following values: mandatory (M), optional (O), conditionally optional (CO), and conditionally mandatory (CM). M indicates that the attribute is mandatory, O indicates that it is optional, and CM indicates that it is conditionally mandatory (i.e., mandatory only if a certain condition is met). CO indicates that it is optional only if a certain condition is met. The "readable" value can take the following values: true (T) and false (F). A value of T indicates that the attribute is readable, while a value of F indicates that it is not readable. The "writable" value also takes the following values: true (T) indicates that the attribute is writable, while a value of F indicates that it is not writable. This explanation also applies to the following tables and will not be repeated here.
[0109] 2.MLEntityLoadingPolicy< <ioc>>
[0110] MLEntityLoadingPolicy< <ioc>> represents the ML model loading policy set by the consumer for the producer. The consumer uses this IOC to set the conditions for triggering ML model loading for the producer. The producer triggers model loading only when the ML model meets the conditions. The properties of this IOC are shown in Table 2.
[0111] Table 2
[0112] The inferenceType attribute indicates the inference type of the initial model associated with the ML model loading policy set by the consumer for the producer. This attribute is conditionally required, meaning that it is required only when the model associated with the ML model loading policy set by the consumer for the producer is the initial training model. It should be understood that the initial training model is the model being trained for the first time and does not yet have a model identifier. Therefore, the inference type of the initial training model can be used to identify the model.
[0113] The mLEntityId attribute is used to indicate the identifier of the retrained model associated with the ML model loading policy set by the consumer for the producer. This attribute is conditionally required. That is, when the model associated with the ML model loading policy set by the consumer for the producer is a retrained model, this attribute is required. It should be understood that an mLEntityId uniquely identifies a model that has already been initially trained.
[0114] The policyForLoading attribute is used to indicate the ML model loading policy set by the consumer to the producer. The policy can be a series of thresholds. For example, the policy can be that the accuracy of the ML model is greater than 90%, or a preset time point or preset period is reached, or a certain network performance indicator is lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule).
[0115] 3. MLEntityLoadingProcess <ioc>>
[0116] MLEntityLoadingProcess< <ioc>> represents an ML model loading process. In the scenario described in sub-use case 1 above, the producer can use this IOC to instantiate one or more ML model loading processes for each ML model loading request from the consumer. In the scenario described in sub-use case 2 above, the producer can use this IOC to instantiate one or more ML model loading processes and associate these one or more ML model loading processes with the ML model loading policy set by the consumer for the producer. The properties of this IOC are shown in Table 3.
[0117] Table 3
[0118] The progressStatus attribute indicates the status of the ML model loading process. The value of this attribute can be any of the following: RUNNING, CANCELLING, SUSPENDED, FINISHED, or CANCELLED. The cancelProcess attribute indicates whether the consumer cancels the ML entity loading process. The value is true (TRUE) or false (FALSE). The suspendProcess attribute indicates whether the consumer suspends the ML entity loading process. The value is true (TRUE) or false (FALSE). The resumeProcess attribute indicates whether the consumer resumes the ML entity loading process. The value is true (TRUE) or false (FALSE).
[0119] The role-related attribute MLEntityLoadingRequestRef is the loading request identifier associated with the ML entity loading process and is required when the ML entity loading process corresponds to the scenario described in Sub-Use Case 1. MLEntityLoadingPolicyRef is the loading policy identifier associated with the ML entity loading process and is required when the ML entity loading process corresponds to the scenario described in Sub-Use Case 2. LoadedMLEntityRef is the model identifier associated with the ML entity loading process.
[0120] Figure 2 illustrates a diagram of the network resource model (NRM) relationships for the IOCs corresponding to Tables 1, 2, and 3. As shown in Figure 2, AiMlInferenceFunction represents a function that can execute ML inference and is also an IOC. In the figure, the lines connecting the MLEntityLoadingRequest, MLEntityLoadingProcess, and MLEntityLoadingPolicy objects with the AiMlInferenceFunction object are solid diamonds on the AiMlInferenceFunction object side. These solid diamonds represent inclusion relationships, meaning that the AiMlInferenceFunction object includes the MLEntityLoadingRequest, MLEntityLoadingProcess, and MLEntityLoadingPolicy objects. Furthermore, the lines connecting the MLEntityLoadingRequest object, the MLEntityLoadingProcess object, and the MLEntity object are arrows on the MLEntity object side. These arrows represent the association relationship, meaning that the MLEntityLoadingRequest object and the MLEntityLoadingProcess object are each associated with the MLEntity object. Therefore, this can be further understood as the AiMlInferenceFunction object containing the MLEntity object. In the figure, the lines connecting the AiMlInferenceFunction object and the MLEntity object are solid diamonds on the AiMlInferenceFunction object side.
[0121] In addition, the numbers and symbols on the lines connecting the various objects in the diagram represent the cardinality relationships between the intent object attributes that have inclusion and association relationships. For example, in the line between the MLEntityLoadingRequest object and the AiMlInferenceFunction object shown in Figure 2, the cardinality on the MLEntityLoadingRequest object side is "*", and the cardinality on the AiMlInferenceFunction object side is "1". On the line between the MLEntityLoadingRequest object and the MLEntity object, the cardinality on both the MLEntityLoadingRequest object and the MLEntity object is "1". This means that an AiMlInferenceFunction object can contain multiple MLEntityLoadingRequest objects, while each MLEntityLoadingRequest object can only be associated with one MLEntity object. The other lines in Figure 2 and the connection rules of the subsequent NRM diagrams can also be interpreted according to this rule and will not be repeated here.
[0122] As shown in Figure 2, the MLEntityLoadingRequest, MLEntityLoadingProcess, and MLEntityLoadingPolicy objects are contained within the AiMlInferenceFunction object. For consumers, each MLEntityLoadingRequest or MLEntityLoadingPolicy created during model deployment targets a single inference function (e.g., the software or entity that performs the inference function). This means that during model deployment, the MLEntityLoadingRequest or MLEntityLoadingPolicy explicitly or implicitly indicates only one target inference function. Therefore, producers can only deploy models for this target inference function. If models need to be deployed for multiple inference functions, consumers must maintain a network resource model relationship, as shown in Figure 2, for each inference function. This requires creating an MLEntityLoadingRequest or MLEntityLoadingPolicy for each inference function to implement model deployment, resulting in a complex and inefficient process.
[0123] In view of this, the present application provides a model deployment method and a communication device, which indicate the deployment range of the target model during the model deployment process, so that one model deployment process can correspond to multiple entities within the deployment range that need to use the target model for reasoning, reducing the complexity of deploying target models for multiple entities and improving the model deployment efficiency.
[0124] It should be understood that in the embodiment of the present application, the inference function (AiMlInferenceFunction) can be understood as an entity that undertakes the inference function, such as a base station. To avoid ambiguity, the inference function will be represented by "entity" in the subsequent description and will not be repeated.
[0125] Below, the design for IOC proposed in this application is first described in detail.
[0126] 1. Add the following properties to the model loading request object created by the consumer:
[0127] Attribute 1: Used to indicate the deployment scope of the target model corresponding to the model loading request created by the consumer. The deployment scope can be used to indicate multiple entities that need to use the target model for reasoning;
[0128] Attribute 2: used to indicate the management method during the deployment of the target model, for example, batch management or independent management;
[0129] Attribute 3: Used to indicate the entity that the consumer specifies that actually needs to deploy the target model.
[0130] For example, a model loading request created by a consumer can be generated through MLEntityLoadingRequest< <ioc>> indicates that Table 4 exemplarily shows an MLEntityLoadingRequest< provided in an embodiment of the present application <ioc>>Contained attributes.
[0131] Table 4
[0132] In one possible implementation, the above-mentioned attribute 1 can be a role-related attribute aIMlInferenceFunctionRef, which can be a conditional optional attribute. For example, when the model loading request created by the consumer is for an entity within a certain range, the attribute is optional. For example, when the deployment range is a geographical area, optionally, a specific list of inference functions within the geographical area can be further specified, thereby selecting a part of the inference functions specified within the geographical area, or it can be not specified, in which case all the inference functions within the geographical area will be selected; the above-mentioned attribute 2 can be infFuncBatchIndicator, which is an optional attribute. The value of infFuncBatchIndicator can indicate the management method in the model deployment process, such as batch management of the model deployment process of multiple entities, or independent management of the model deployment process of each entity. It should be understood that the embodiment of the present application does not specifically limit the expression form of attribute 1 and attribute 2.
[0133] In one possible implementation, the above-mentioned attribute 3 can be a role-related attribute aIMlInferenceFunctionToLoadMlEntityRef, which can be a conditionally optional attribute. For example, when the model loading request created by the consumer is used to indicate the deployment of the target model for some inference functions within the above-mentioned deployment scope, this attribute is optional. The value of this attribute can represent the identifier of the part of the inference function specified by the consumer, but this application does not make any specific limitations on this.
[0134] 2. Add the following properties to the model loading strategy object set by the consumer to the producer:
[0135] Attribute 4: Used to indicate the deployment scope of the target model corresponding to the model loading strategy set by the consumer to the producer. This deployment scope can be used to indicate multiple entities that need to use the target model for reasoning;
[0136] Attribute 5: Used to indicate the entity that the consumer specifies that actually needs to deploy the target model.
[0137] For example, the model loading policy set by the consumer to the producer can be set by MLEntityLoadingPolicy< <ioc>> indicates that Table 5 exemplarily shows an MLEntityLoadingPolicy provided in an embodiment of the present application. <ioc>>Contained attributes.
[0138] Table 5
[0139] In one possible implementation, attribute 4 may be the role-related attribute aIMlInferenceFunctionRef. This attribute may be conditionally optional. For example, when the model loading policy set by the consumer for the producer targets entities within a certain range, this attribute may be optional. It should be understood that attribute 4 may also have other forms, which are not limited in this application.
[0140] In one possible implementation, the above-mentioned attribute 5 can be IMlInferenceFunctionToLoadMlEntityRef, which can be a conditionally optional attribute. For example, when the model loading strategy set by the consumer to the producer includes deploying the target model for some inference functions within the deployment scope of the target model, this attribute is optional. The value of this attribute can represent the identifier of the part of the inference function specified by the consumer, but this application is not limited to this.
[0141] In one possible implementation, the model loading process object can be accessed through MLEntityLoadingProcess< <ioc>> indicates that the attributes it contains may be as shown in Table 3, but this application does not limit this.
[0142] Figure 3 illustrates a network resource model relationship diagram provided by an embodiment of the present application. Compared to Figure 2, the MLEntityLoadingRequest object, MLEntityLoadingProcess object, and MLEntityLoadingPolicy object are no longer contained in an AiMlInferenceFunction object, but are instead contained in a ManagedEntity object, which is a proxy class (ProxyClass).
[0143] Exemplarily, ManagedEntity can represent one or more of the following: a group of managed entities, which can be represented by a network area, such as a subnet (SubNetwork), which can include a group of base stations; a group of managed functions (ManagedFunction), such as a communication function implemented by software running on dedicated hardware, or a communication function implemented by software running on a network function virtualization infrastructure (NFVI), which can be a virtual network function of a base station or a core network; or a group of managed management functions (ManagementFunction), such as an inference function (AI / ML inference function), which can also be understood as the bearer of the inference function.
[0144] In addition, on the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityLoadingPolicy object, and on the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityLoadingRequest object, the cardinality close to the MLEntityLoadingProcess object has been changed from "*" (shown in Figure 2) to "1..*", indicating that at least one MLEntityLoadingProcess object can be associated with one MLEntityLoadingPolicy object, and at least one MLEntityLoadingProcess object can be associated with one MLEntityLoadingRequest object.
[0145] It should be understood that based on the network resource model shown in FIG3 , an MLEntityLoadingRequest or an MLEntityLoadingPolicy created by a consumer during a model deployment process can no longer target only one entity (such as an AiMlInferenceFunction shown in FIG2 ), but can target multiple entities (such as the ManagedEntity shown in FIG3 ).
[0146] Figure 4 is a schematic diagram of a communication system 400 applicable to an embodiment of the present application. As shown in Figure 4, the communication system 400 includes a first device 410, a second device 420, and a third device 430. The first device 410, the second device 420, and the third device 430 can establish communication with each other.
[0147] In one possible implementation, the role of the first device 410 can be a model loading consumer (ML loading consumer) that can call the model loading service, the role of the second device 420 can be a model loading producer (ML loading producer) that can provide model loading services, and the third device 430 can be a role that implements the inference function (ML inference) based on the model (it can also be understood as the bearer of the inference function, the executor of the inference function, and the obtainer of the inference result. In some implementations, this role can be directly called the inference function), but this application does not make any specific limitations on this.
[0148] It should be understood that in the communication system shown in FIG4 , the number of the first device and the second device can be one or more, and the number of the third device can be multiple, and this application does not limit this.
[0149] In one possible scenario, the first device 410 may call the model loading service provided by the second device 420 to instruct the target model to be loaded into the designated third device 430 , and the third device 430 is responsible for inputting data into the model to obtain corresponding output.
[0150] In one possible implementation, the first device 410 (or the second device 420, or the third device 430) may include one or more of the following: a module that can implement the corresponding function of model loading consumers, a module that can implement the corresponding function of model loading producers, or a module that can implement reasoning functions based on the model. This application does not make specific limitations on this.
[0151] Optionally, the first device 410 can be a network management system (NMS), an element management system (EMS), or a radio access network (RAN) device, or any other device or system that can be used as a model to load consumers, and this application does not make specific limitations on this.
[0152] Optionally, the second device 420 may be a network management system, a network element management system, or a wireless access network device, or any other device or system that can serve as a model loading producer, and this application does not impose any specific limitations on this.
[0153] Optionally, the third device 430 can be a network management system, a network element management system, or a wireless access network device, or any other device or system that can implement reasoning functions based on a model, and this application does not make any specific limitations on this.
[0154] Optionally, the above-mentioned network element management system can manage the network and be responsible for the operation, management and functional maintenance of the network. It can also be called a cross-domain management system, which is not limited in this application.
[0155] Optionally, the network element management system can be used to manage one or more network elements of a certain category, and can also be called a domain management system or a single domain management system, which is not limited in this application.
[0156] Optionally, the above-mentioned access network device can be a transmission reception point (TRP), an evolved NodeB (eNB or eNodeB) in a long term evolution (LTE) system, a home base station (for example, home evolved NodeB, or home Node B, HNB), a base band unit (BBU), or a wireless controller in a cloud radio access network (CRAN) scenario, or the access network device can be a relay station, an access point, a vehicle-mounted device, a wearable device, and an access network device in a 5G network or an access network device in a future evolved public land mobile communication network (PLMN) network, etc., it can be an access point (AP) in a wireless local area network (WLAN), it can be a gNB in a new radio (NR) system, it can be a satellite base station in a satellite communication system, etc., and the embodiments of the present application are not limited.
[0157] Optionally, the access network device may include a centralized unit (CU) node, a distributed unit (DU) node, or an access network device including a CU node and a DU node, or an access network device including a control plane CU node (CU-CP node) and a user plane CU node (CU-UP node) and a DU node. The access network device including the CU node and the DU node may split the protocol layer of the access network device, with the functions of some protocol layers being centrally controlled by the CU, and the functions of the remaining part or all of the protocol layers being distributed in the DU, which is centrally controlled by the CU. As an implementation method, the CU deployment protocol stack includes a radio resource control (RRC) layer, a packet data convergence protocol (PDCP) layer, and a service data adaptation protocol (SDAP) layer. The DU deployment protocol stack includes a radio link control (RLC) layer, a media access control (MAC) layer, and a physical layer (PHY) layer. Thus, the CU has the processing capabilities of RRC, PDCP, and SDAP. DU has the processing capabilities of RLC, MAC and PHY. The above-mentioned functional division is only an example and does not constitute a limitation on CU and DU. That is to say, there may be other ways of functional division between CU and DU, which will not be described in detail in the embodiments of the present application. The functions of CU can be implemented by one entity or by different entities. For example, the functions of CU can be further divided, for example, the control plane (CP) and the user plane (UP) are separated, that is, the control plane of CU (CU-CP) and the user plane of CU (CU-UP). For example, CU-CP and CU-UP can be implemented by different functional entities, and CU-CP and CU-UP can be coupled with DU to jointly complete the functions of access network equipment. In one possible way, CU-CP is responsible for control plane functions, mainly including RRC and PDCP-C, where PDCP-C is mainly responsible for encryption and decryption, integrity protection, data transmission, etc. of control plane data. CU-UP is responsible for user plane functions, mainly including SDAP and PDCP-U, where SDAP is mainly responsible for processing data of core network devices and mapping data flows to bearers. The PDCP-U is responsible for data plane encryption and decryption, integrity protection, header compression, sequence number maintenance, and data transmission. The CU-CP and CU-UP are connected via the E1 interface. The CU-CP represents access network equipment and connects to core network equipment via the interface between the core network and access network equipment.The DU is connected via F1-C (control plane). The CU-UP is connected to the DU via F1-U (user plane). In addition, another possible implementation is that the PDCP-C is also in the CU-UP, which is not limited in this application.
[0158] In order to make the purpose and technical solution of this application clearer and more intuitive, the model deployment method and communication device provided by the embodiment of this application will be described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0159] Below, a model deployment method provided in an embodiment of the present application is first described in detail with reference to Figures 5 to 10.
[0160] The model deployment method provided in the embodiments of this application can be implemented based on the design for IOC provided in this application. This method can be applied to the communication system 400 shown in Figure 4 or other communication systems, and this application does not limit this. For example, in the embodiments of this application, the model deployment method provided in this application is described from the perspective of inter-device interaction, taking the role of the first device as a consumer and the role of the second device as a producer as an example.
[0161] FIG5 is a schematic flow chart of a model deployment method 500 provided in an embodiment of the present application. The method 500 includes the following steps:
[0162] S501. The first device determines identification information of a target model and a deployment range of the target model. The deployment range of the target model is used to indicate a plurality of entities that need to use the target model for reasoning.
[0163] It should be noted that the model in this application can be a machine learning model (ML model) or can be embodied in the form of an entity, such as a machine learning entity (ML entity), and this application does not limit this.
[0164] S502: The first device sends identification information of the target model and the deployment range of the target model. Correspondingly, the second device receives the identification information of the target model and the deployment range of the target model.
[0165] S503: The second device deploys the target model based on the identification information of the target model and the deployment range of the target model.
[0166] It should be noted that deploying the target model by the second device means that the second device loads the target model to one or more target reasoning functions corresponding to the deployment scope, where the target model is used to perform reasoning.
[0167] In an embodiment of the present application, the information sent by the first device to the second device includes the deployment scope of the target model, which is used to indicate multiple entities that need to use the target model for reasoning. The second device can deploy the target model for all or part of the multiple entities indicated by the deployment scope based on the deployment scope sent by the first device. Compared with some examples, in scenarios where models need to be deployed for multiple entities, the first device needs to send identification information of the target model to the second device multiple times for multiple entities. The method provided in the present application indicates the deployment scope of the target model during the model deployment process, so that one model deployment process can correspond to multiple entities within the deployment scope that need to use the target model for reasoning, reducing the complexity of deploying target models for multiple entities and improving the efficiency of model deployment.
[0168] As an optional embodiment, the deployment scope of the target model in the above S501 can be indicated by one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple reasoning functions; or identification information of multiple entities.
[0169] It should be understood that the deployment scope of the target model may refer to a scope that includes multiple entities that need to use the target model for reasoning, which may be a substantial regional scope or an abstract scope, and this application does not limit this.
[0170] Optionally, the information of the geographical area may be administrative division information of the geographical area, longitude and latitude information of the geographical area, etc., which is not limited in this application. It is understandable that a first mapping relationship may exist between the information of the geographical area and the entity that performs the reasoning function, and the first device and / or the second device may maintain the first mapping relationship. In one possible implementation, the second device may determine multiple entities in the geographical area as entities that need to be reasoned using the target model based on the information of the geographical area indicated by the first device and the first mapping relationship. Furthermore, the second device may perform subsequent steps for the determined multiple entities.
[0171] Optionally, the identification information of the network area can be the identification information of any subnet that can distinguish the network area, such as a local area network, a metropolitan area network, a wide area network, etc., and this application does not limit this. It is understandable that a second mapping relationship can exist between the identification information of the network area and the entity that performs the reasoning function, and the first device and / or the second device can maintain the second mapping relationship. In one possible implementation, the second device can determine multiple entities that need to be reasoned using the target model based on the identification information of the network area indicated by the first device and the second mapping relationship, and further perform subsequent steps on the determined multiple entities.
[0172] In some examples, an entity that performs an inference function may be referred to as an inference function (AIM1InferenceFunction). Optionally, identification information of multiple inference functions may be used to identify multiple entities that perform inference functions. For example, the deployment scope of the target model may be indicated by multiple AIM1InferenceFunction identifiers or an AIM1InferenceFunction list.
[0173] Optionally, the identification information of the multiple entities can be any information that can uniquely identify the entities, such as the Internet Protocol (IP) addresses, Media Access Control (MAC) addresses, etc. of the multiple entities, and this application does not limit this. In some examples, the identification information of the multiple entities can be presented in the form of a list, but this application does not specifically limit this.
[0174] As an optional embodiment, a possible implementation of the above S503 includes: the second device sends first information to the target entity, where the first information is used to instruct the target entity to deploy the target model.
[0175] It should be understood that the deployment scope of the above-mentioned target model indicates multiple entities that need to use the target model for reasoning. In the embodiment of the present application, the target entity can be understood as the entity that actually deploys the target model. The target entity may include all or part of the multiple entities indicated by the deployment scope of the above-mentioned target model.
[0176] In a possible implementation, the target entity is determined by the second device from a deployment scope of the target model, and the target entity is determined by the second device based on resources and / or capabilities of multiple entities.
[0177] In another possible implementation, the target entity is specified by the first device. Before the second device sends the first information to the target entity, the method further includes: the first device sending identification information of the target entity to the second device, where the target entity is determined by the first device based on resources and / or capabilities of multiple entities. Correspondingly, the second device receives the identification information of the target entity.
[0178] In a possible scenario, the target entity includes all entities in the multiple entities, and the first information includes identification information of the target model.
[0179] In this scenario, the multiple entities indicated by the deployment scope of the target model are all target entities. Each of the multiple entities can use the target model to perform reasoning by performing reasoning based on its deployed target model and obtaining a reasoning result. In this case, the first information sent by the second device to the target entity includes identification information of the target model, thereby instructing the target entity to deploy the target model.
[0180] In some examples, the target entity already has a target model stored therein, and the second device sends first information to the target entity, the first information including identification information of the target model, instructing the target entity to deploy the target model. The second device instructing the target entity to deploy the target model can be understood as the second device instructing the target entity to activate the target model.
[0181] In other examples, if the target entity does not store the target model, the second device may send the first information to the target entity to instruct the target entity to deploy the target model. This can be understood as the second device instructing the target entity to download and activate the target model. In some implementations, the first information also includes the storage address of the target model, and the target entity downloads the target model based on the storage address. In other implementations, the target entity maintains a correspondence between the identification information of the target model and the storage address of the target model. The target entity can search and obtain the storage address based on the identification information of the target model included in the first information, and download the target model based on the storage address.
[0182] In another possible scenario, the target entity includes some entities among the multiple entities, and the first information includes identification information of the target model and a deployment range of the target model.
[0183] In this scenario, the multiple entities indicated by the deployment scope of the target model include the target entity and other entities except the target entity. In this case, the target entity can use the target model deployed by itself for reasoning, while other entities can use the target model deployed on the target entity for reasoning.
[0184] In one possible implementation, taking the first entity as an example, other entities may utilize a target model deployed on a target entity for inference, including: the first entity sending a first request to the target entity, the first request being used to request an inference result from the target model, the first request including identification information of the target model and input parameters of the target model; the target entity correspondingly receiving the first request and, upon determining that the first entity falls within the deployment scope of the target model, obtaining an inference result using the target model deployed on the target entity based on the input parameters. Furthermore, the target entity sends the inference result to the first entity, and the first entity correspondingly receives the inference result.
[0185] In one possible implementation, before the first entity sends the first request to the target entity, the method further includes: a second device sending second information to the first entity, the second information being used to indicate successful deployment of the target model, the second information including identification information of the target model and identification information of the target entity. Accordingly, the first entity receives the second information.
[0186] It should be understood that after the second device successfully deploys the target model on the target entity, it can notify the first entity of the successful deployment of the target model through the second information. The identification information of the target model in the second information is used to indicate the successfully deployed model, and the identification information of the target entity is used to indicate on which entity the target model is deployed.
[0187] As an optional embodiment, after S502 above, method 500 further includes: the second device sending first identification information and identification information of the first process, where the first identification information is associated with a message carrying identification information of the target model, the first identification information and the first process are mapped to each other, and the first process is used to manage the deployment process of the target model. Correspondingly, the first device receives the first identification information and the identification information of the first process.
[0188] Optionally, the first identification information and the identification information of the first process may be sent via a response message, and the response message may be a response to a message carrying the identification information of the target model, but this application does not limit this.
[0189] It should be understood that the embodiments of the present application are applicable to the scenarios corresponding to the two sub-use cases defined in the model deployment phase mentioned above. In different scenarios, the identification information of the target model can be carried by different types of messages, so the meaning of the first identification information is different in different scenarios.
[0190] Scenario 1: Scenario corresponding to the above-mentioned sub-use case 1. In a possible implementation of this scenario, the identification information of the target model in the above S501 is sent by the first device through a request message.
[0191] In a possible implementation, the request message is used to request the second apparatus to deploy the target model for multiple entities indicated by the deployment scope of the target model, which may correspond to the case described above where the target entity includes all entities in the multiple entities.
[0192] In another possible implementation, the request message is used to request the second device to deploy the target model for some of the multiple entities indicated by the deployment scope of the target model, which may correspond to the situation described above where the target entity includes some of the multiple entities.
[0193] Optionally, the deployment scope of the target model sent by the first device to the second device in the above S501 can also be sent through a request message. The deployment scope of the target model and the identification information of the target model can be sent through the same request message or through different request messages. This application does not limit this.
[0194] It should be understood that in this scenario, the first device determines the timing of sending the above-mentioned request message to the second device based on the deployment strategy of the target model. The deployment strategy can be the requirement for the reasoning ability of the target model, or the parameters of the various functional entities involved in the model deployment process, etc. The deployment strategy can be determined by the first device, or it can be reported by the entity that needs to use the target model for reasoning. This application does not limit this. For example, the deployment strategy can be that the accuracy of the target model is greater than 90%, or the time point or preset period for deploying the target model is reached in time, or a certain network performance indicator of the entity is monitored to be lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations on this.
[0195] In one possible implementation, if a first device determines that a target model satisfies conditions corresponding to a deployment policy and / or that all or some entities indicated by a deployment scope of the target model satisfy the deployment conditions, the first device sends a request message to a second device requesting deployment of the target model. The request message includes identification information of the target model and the deployment scope of the target model. In response, the second device receives the request message and creates an instance of the request message and a first process instance.
[0196] It should be understood that the first process is used to manage the deployment process of the target model. The first process can be referred to as the target model deployment process or the target model loading process. This application does not specifically limit the specific name of the first process. In addition, unless otherwise specified, the terms "loading" and "deployment" described in the embodiments of this application can be understood to have the same meaning and will not be repeated hereafter.
[0197] Illustratively, in this scenario, the request message may correspond to the MLEntityLoadingRequest object designed in this application, and its properties may be as shown in Table 4.
[0198] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingRequest object, and the identification information of the target model corresponding to the request message can be represented by the value of the role-related attribute mLEntityToLoadRef of the MLEntityLoadingRequest object, but this application does not make specific limitations on this.
[0199] Correspondingly, the first process may be a process created based on the MLEntityLoadingProcess object shown in Table 3. The request message may be represented by the value of the role-related attribute MLEntityLoadingRequestRef of the MLEntityLoadingProcess object, and the identification information of the target model associated with the first process may be represented by the value of the role-related attribute LoadedMLEntityRef.
[0200] In the above scenario 1, the instance of the request message may be an MLEntityLoadingRequest instance, the first identification information may be the identifier of the MLEntityLoadingRequest instance, the instance of the first process may be an MLEntityLoadingProcess instance, and the identification information of the first process may be the identifier of the MLEntityLoadingProcess instance.
[0201] Scenario 2: Scenario corresponding to the above-mentioned sub-use case 2. In a possible implementation of this scenario, the identification information of the target model in the above S501 is sent by the first device through a configuration message, and the configuration message also includes a deployment strategy of the target model.
[0202] In one possible implementation, the first device sends the identification information of the target model and the deployment policy of the target model to the second device through a configuration message. After receiving the identification information of the target model and the deployment policy of the target model, the second device creates an instance of the configuration message, and after determining that the target model has met the conditions corresponding to the deployment policy, and / or that all or part of the entities indicated by the deployment scope of the target model have met the deployment conditions, creates an instance of the first process.
[0203] Optionally, the deployment scope of the target model sent by the first device to the second device in the above S501 can also be sent through a configuration message. The deployment scope of the target model, the identification information of the target model and the deployment strategy of the target model can be sent through the same configuration message or through multiple configuration messages. This application does not limit the number, content and sequence of configuration messages.
[0204] Optionally, the deployment strategy of the target model in this scenario may be that the accuracy of the target model is greater than 90%, or that the preset time point or preset period for deploying the target model is reached, or that a certain network performance indicator of the third device is monitored to be lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations on this.
[0205] Illustratively, in this scenario, the configuration message may correspond to the MLEntityLoadingPolicy object designed in this application, and its properties may be as shown in Table 5.
[0206] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingPolicy object, the deployment policy of the target model can be represented by the value of the attribute policyForLoading, and the identification information of the target model can also be represented by the value of the attribute inferenceType or the value of the attribute mLEntityId, but this application does not limit this.
[0207] Correspondingly, the first process may be a process created based on the MLEntityLoadingProcess object described in Table 3. The configuration message may be identified by the value of the role-related attribute MLEntityLoadingPolicyRef of the MLEntityLoadingProcess object, and the identification information of the target model associated with the first process may be represented by the value of the role-related attribute LoadedMLEntityRef.
[0208] In the above scenario 2, the instance of the configuration message may be an MLEntityLoadingPolicy instance, the first identification information may be the identifier of the MLEntityLoadingPolicy instance, the instance of the first process may be an MLEntityLoadingProcess instance, and the identification information of the first process may be the identifier of the MLEntityLoadingProcess instance.
[0209] In both scenarios described above, the first process is used to manage the target model's deployment process. If there are multiple target entities to which the target model needs to be deployed, the target model deployment process management methods can include "batch management" and "independent management," but this application does not specifically limit these methods. The meaning of the first process varies under different deployment process management methods. The following uses "batch management" and "independent management" as examples for detailed explanation.
[0210] The first method: batch management
[0211] "Batch management" can be understood as treating the target model deployment process for multiple target entities as a whole. In this case, the first process corresponds to the entire target model deployment process. For example, the second device creates an MLEntityLoadingProcess instance to manage the target model deployment process for multiple target entities.
[0212] It should be understood that the MLEntityLoadingProcess instance has a status attribute, progressStatus, which can take any of the following values: RUNNING, CANCELLING, SUSPENDED, FINISHED, or CANCELLED. When the progressStatus value of an MLentityLoadingProcess instance is RUNNING, it can be understood that the model deployment process is in progress. Therefore, RUNNING can also be understood as "deploying" or "loading", which is not limited in this application.
[0213] Exemplarily, the second apparatus may send the first information to the target entity while creating the MLEntityLoadingProcess instance, or after creating the MLEntityLoadingProcess instance when the progressStatus value of the MLEntityLoadingProcess instance is RUNNING.
[0214] Optionally, the second device may report the status of the first process to the first device.
[0215] In the case where the deployment process management mode is batch management, the number of the first process may be one, and the first process corresponds to the overall process of the deployment model of multiple target entities.
[0216] In one possible implementation, the second device may proactively report the status of the first process to the first device. The method further includes: the second device sending third information, the third information including the first identification information and the status of the first process, or the identification information of the first process and the status of the first process. Correspondingly, the first device receives the third information.
[0217] It should be understood that the third information is used to report the status of the first process. The first identification information is associated with the message carrying the identification information of the target model and has a mapping relationship with the first process. Therefore, the second device can use the first identification information or the identification information of the first process to identify the first process. The status of the first process can include: running, canceling, canceled, paused, or ended.
[0218] In another possible implementation, the second device may report the status of the first process to the first device in response to the second request from the first device. Specifically, before the second device sends the third information, the second device receives the second request from the first device, where the second request is used to request the status of the first process, and the first request includes the first identification information, or the identification information of the first process.
[0219] Optionally, the first device may control the first process. In one possible implementation, the method further includes: the first device sending fourth information, the fourth information being used to control the overall process of deploying the target model by multiple target entities, the fourth information including the first identification information and a control instruction for the first process, or the identification information of the first process and a control instruction, the control instruction including canceling, pausing, or starting. Correspondingly, the second device receives the fourth information.
[0220] In an embodiment of the present application, through batch management of model deployment, the number of first processes can be one, and the second device can maintain only one MLEntityLoadingProcess instance, thereby achieving the purpose of deploying the target model to multiple target entities, thereby facilitating the second device's management of the model deployment process. Furthermore, during the interaction between the first and second devices regarding status requests and reports of the first process and control of the first process, the process of deploying the target model to multiple target entities is treated as a whole. There is no need to monitor the model deployment process for each target entity, nor is there a need to exchange information between the target entities. This helps reduce signaling overhead between devices and improves the efficiency of target model deployment.
[0221] The second method: independent management
[0222] "Independent management" can be understood as creating a separate process to manage the model deployment process for each of multiple target entities. In this case, the first process includes multiple processes, each of which corresponds to the model deployment process of a target entity. The identification information of the first process includes multiple identification information corresponding to the multiple processes. Exemplarily, the second device creates multiple MLEntityLoadingProcess instances to manage the target model deployment process. These multiple MLEntityLoadingProcess instances correspond to the model deployment processes of the target entities.
[0223] When the deployment process management method is independent management, when the second device reports the status of the first process to the first device, the target model deployment process corresponding to this entity can be identified by the first identification information and the identification information of a certain entity in the target entity, or the target model deployment process of a target entity can be identified by the identification information of a certain process in multiple processes.
[0224] In one possible implementation, the second device actively sends fifth information, which is used to report the status of all or part of the multiple processes. The fifth information includes first identification information, identification information of each entity corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, identification information of all or part of the processes, and the respective status of all or part of the processes.
[0225] In another possible implementation, the second device may report the status of all or some of the multiple processes to the first device in response to a third request from the first device. Specifically, before the second device sends the fifth information, the second device receives a third request from the first device, where the third request is used to request the status of all or some of the multiple processes, and the third request includes the first identification information, identification information of each entity corresponding to all or some of the multiple processes, or identification information of all or some of the processes.
[0226] Optionally, the first device may control multiple processes included in the first process based on the control instruction. In one possible implementation, the method further includes: the first device sending sixth information, the sixth information being used to control all or some of the multiple processes, the sixth information including the first identification information, identification information of entities corresponding to all or some of the processes, and control instructions for each of the all or some of the processes; or identification information of all or some of the processes and the control instructions.
[0227] In an embodiment of the present application, through an independently managed model deployment method, the first process may include multiple processes, and the multiple processes correspond one-to-one to the identification information of multiple target entities, and the multiple processes correspond to the target model deployment processes of the multiple target entities. In this way, the first device can respectively obtain and control the status corresponding to the model deployment process of each target entity. In the case that some target entities among the multiple target entities fail to be deployed, it can be determined which target entities failed to be deployed, and based on the identification information of the target entities that failed to be deployed, the target models can be re-deployed for these target entities, which is conducive to improving management accuracy and improving the efficiency of model deployment.
[0228] In one possible implementation, the target model's deployment process management mode is a default mode, which includes batch management or independent management, that is, no explicit instructions are required from either party. Optionally, the deployment process management mode can be agreed upon in the protocol, which is not limited in this application.
[0229] In another possible implementation, the deployment process management mode of the target model is indicated by the first device. In one embodiment, the first device sends first indication information to the second device, where the first indication information is used to indicate the deployment process management mode of the target model.
[0230] Optionally, the first indication information may be represented by the value of the attribute infFuncBatchIndicator. In the above scenario 1, the first indication information may be represented by the value of the attribute infFuncBatchIndicator of the MLEntityLoadingRequest object, but this application does not limit this.
[0231] Optionally, the value of the first indication information infFuncBatchIndicator may include 0 and 1. When the value is 0, it indicates that the deployment process management mode of the target model is batch management. When the value is 1, the deployment process management mode of the target model is independent management. This application does not specifically limit the value of infFuncBatchIndicator and the meaning of the value.
[0232] Below, the two scenarios of scenario 1 and scenario 2, the two deployment process management methods of batch management and independent management, and the two situations where the target entity includes all entities of multiple entities and the target entity includes some entities of multiple entities are described in detail in combination with Figures 6 to 10.
[0233] For ease of understanding, in the descriptions of Figures 6 to 10, the entity is a gNB, the multiple entities include a first gNB, a second gNB and a third gNB, the first device in Figure 5 is an NMS, and the second device is an EMS. It should be understood that the number of entities is only an example and does not constitute a limitation to this application.
[0234] Case 1: Scenario 1 + target entity includes all entities in multiple entities + batch management
[0235] FIG6 is a schematic flow chart of a model deployment method 600 provided in an embodiment of the present application. The method 600 includes:
[0236] S601: When the NMS determines that the deployment policy of the target model is satisfied, it sends an MLEntityLoadingRequest to the EMS. The MLEntityLoadingRequest includes identification information of the target model and the deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingRequest.
[0237] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingRequest object shown in Table 4, and the identification information of the target model can be represented by the value of the role-related attribute mLEntityToLoadRef of the MLEntityLoadingRequest object, but this application does not make specific limitations on this.
[0238] For example, in this embodiment, the deployment scope of the target model indicates that the entities requiring inference using the target model are the first gNB, the second gNB, and the third gNB. It should be understood that the number of entities presented in this embodiment is for illustrative purposes only and does not constitute a limitation of this application. The number of entities involved in this application may be greater or lesser.
[0239] S602: The EMS creates an MLEntityLoadingRequest instance and an MLEntityLoadingProcess instance.
[0240] It should be understood that the MLEntityLoadingProcess instance created by the EMS in S602 corresponds to the model deployment process of the first gNB, the second gNB, and the third gNB, that is, the process of deploying the target model for the first gNB, the second gNB, and the third gNB is regarded as a whole and managed through one process.
[0241] It should also be understood that the target model's deployment process management mode can be either default or NMS-indicated. In the case of NMS-indicated, this indication information can be provided by the infFuncBatchIndicator included in the MLEntityLoadingRequest. In this embodiment, the infFuncBatchIndicator indicates that the target model's deployment process management mode is batch management.
[0242] S603: The EMS sends a response message of the MLEntityLoadingRequest to the NMS, the response message including an identifier of an MLEntityLoadingRequest instance and an identifier of an MLEntityLoadingProcess instance. Correspondingly, the NMS receives the response message.
[0243] S604 : The NMS sends a deployment process status request of the target model to the EMS, including the identifier of the MLEntityLoadingRequest instance. Correspondingly, the EMS receives the request.
[0244] In another implementation of S604 , the target model deployment process status request sent by the NMS to the EMS includes an identifier of the MLEntityLoadingProcess instance.
[0245] S605: The EMS reports the target model deployment process status to the NMS, including the identifier of the MLEntityLoadingRequest instance and the progressStatus value corresponding to the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the target model deployment process status.
[0246] In another implementation of S605 , the EMS reports the deployment process status of the target model to the NMS, including the identifier of the MLEntityLoadingProcess instance and the progressStatus value corresponding to the MLEntityLoadingProcess instance.
[0247] S606: The NMS sends a control message for the target model deployment process to the EMS, including the identifier of the MLEntityLoadingRequest instance and a control instruction. Correspondingly, the EMS executes the control instruction.
[0248] In another implementation of S606 , the NMS sends a control message of the target model deployment process to the EMS, including the identifier of the MLEntityLoadingProcess instance and a control instruction.
[0249] S607: The EMS sends first information to the first gNB, the second gNB, and the third gNB. The first information indicates a target deployment model and includes identification information of the target model. Correspondingly, the first gNB, the second gNB, and the third gNB receive the first information.
[0250] Optionally, the way in which the EMS sends the first information to the first gNB, the second gNB, and the third gNB can be unicast or broadcast, which is not limited in this application.
[0251] It should be understood that S607 can be understood as the process of EMS performing model deployment and loading the target model into the first gNB, the second gNB and the third gNB.
[0252] S608. The first gNB, the second gNB, and the third gNB deploy the target model based on the identification information of the target model.
[0253] It should be understood that the gNB deploying the target model in S608 can be understood as the process of the gNB activating the target model.
[0254] S609: The first gNB, the second gNB, and the third gNB perform inference based on the target model.
[0255] It should be understood that the first gNB, the second gNB, and the third gNB each execute S609 when it is necessary to use the target model for inference.
[0256] It should also be understood that the above S604, S605, and S606 are all optional steps.
[0257] In this embodiment of the present application, by including the deployment scope of the target model in the MLEntityLoadingRequest, it is indicated that this model deployment targets multiple gNBs. Through batch management of model deployment, the EMS only maintains a single MLEntityLoadingProcess instance to deploy the target model to multiple gNBs, facilitating EMS management of the model deployment process. Furthermore, when the NMS requests the EMS for the process status of the MLEntityLoadingProcess instance, controls the progress of the MLEntityLoadingProcess instance, or reports the process status of the MLEntityLoadingProcess instance to the NMS, these interactions only need to be performed for a single instance, facilitating management, reducing signaling overhead between interacting entities, and improving target model deployment efficiency.
[0258] Case 2: Scenario 1 + target entity includes all entities in multiple entities + independent management
[0259] FIG7 is a schematic flow chart of a model deployment method 700 provided in an embodiment of the present application. The method 700 includes:
[0260] S701: When the NMS determines that the deployment policy of the target model is satisfied, it sends an MLEntityLoadingRequest to the EMS. The MLEntityLoadingRequest includes identification information of the target model and the deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingRequest.
[0261] It should be understood that the representation of the identification information of the target model and the deployment range of the target model in method 700 may be the same as that in the above method 600, and will not be repeated here.
[0262] S702: The EMS creates an MLEntityLoadingRequest instance and multiple MLEntityLoadingProcess instances.
[0263] It should be understood that the multiple MLEntityLoadingProcess instances correspond to processes for deploying the target model for the first gNB, the second gNB, and the third gNB, respectively. Exemplarily, the number of the multiple MLEntityLoadingProcess instances is three.
[0264] It should also be understood that the target model's deployment process management mode can be either default or NMS-indicated. In the case of NMS-indicated, this indication information can be provided by the infFuncBatchIndicator included in the MLEntityLoadingRequest. In this embodiment, the infFuncBatchIndicator indicates that the target model's deployment process management mode is independent management.
[0265] S703: The EMS sends a response message of MLEntityLoadingRequest to the NMS, where the response message includes the identifier of the MLEntityLoadingRequest instance and multiple identifiers corresponding to multiple MLEntityLoadingProcess instances. Correspondingly, the NMS receives the response message.
[0266] S704. The NMS sends a status request for the MLEntityLoadingProcess instances corresponding to all or some of the gNBs in the multiple gNBs to the EMS, including the identifier of the MLEntityLoadingRequest instance and the identifiers of all or some of the requested gNBs. Correspondingly, the EMS receives the request.
[0267] In another implementation of S704, the NMS sends a status request of the MLEntityLoadingProcess instances corresponding to all or part of the multiple gNBs to the EMS, including the identifiers of the MLEntityLoadingProcess instances corresponding to all or part of the multiple gNBs.
[0268] S705. The EMS reports the status of the MLEntityLoadingProcess instances corresponding to all or some of the gNBs to the NMS, including the identifier of the MLEntityLoadingRequest instance, the identifier of all or some of the gNBs, and the progressStatus value of the MLEntityLoadingProcess instances corresponding to all or some of the gNBs. Correspondingly, the NMS receives the status of the MLEntityLoadingProcess instances corresponding to all or some of the gNBs.
[0269] In another implementation of S705, the EMS reports to the NMS the status of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs, including the identifiers of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs and the progressStatus values of the MLEntityLoadingProcess instances corresponding to all or part of the target gNBs.
[0270] S706. The NMS sends a control message to the EMS regarding the target model deployment process for all or some of the multiple target gNBs. The message includes the identifier of the MLEntityLoadingRequest instance, the identifiers of all or some of the gNBs to be controlled, and a control instruction for the MLEntityLoadingProcess instance corresponding to each gNB. The EMS executes the control instruction accordingly.
[0271] In another implementation of S706, the NMS sends a control message for the target model deployment process of all or part of the gNBs in the multiple gNBs to the EMS, including the identifiers of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs + a control instruction for the MLEntityLoadingProcess instance corresponding to each gNB.
[0272] S707: The EMS sends first information to each of the first gNB, the second gNB, and the third gNB. The first information indicates a target deployment model and includes identification information of the target model. Correspondingly, the first gNB, the second gNB, and the third gNB receive the first information.
[0273] Optionally, the EMS may send the first information to the first gNB, the second gNB, and the third gNB by unicast or broadcast, which is not limited in this application.
[0274] S708. The first gNB, the second gNB, and the third gNB deploy the target model based on the identification information of the target model.
[0275] S709: The first gNB, the second gNB, and the third gNB perform inference based on the target model.
[0276] It should be understood that the above S704, S705, and S706 are all optional steps.
[0277] In this embodiment of the present application, the target model deployment process for multiple gNBs is independently managed. The NMS can obtain and control the process status corresponding to the target model deployment of all or some of the multiple gNBs. If the model deployment fails for some of the multiple gNBs, the NMS can determine which gNBs failed to be deployed and re-deploy the target model for these gNBs, which is conducive to improving management accuracy.
[0278] Case 3: Scenario 1 + target entity includes some entities among multiple entities and the number of some entities is one
[0279] Taking the target model deployed on the first gNB as an example, FIG8 is a schematic flowchart of a model deployment method 800 provided in an embodiment of the present application. The method 800 includes:
[0280] S801: Upon determining that the deployment policy of the target model is satisfied, the NMS sends an MLEntityLoadingRequest to the EMS. The MLEntityLoadingRequest includes identification information of the target model, identification information of the first gNB, and the deployment range of the target model. In response, the EMS receives the MLEntityLoadingRequest.
[0281] Optionally, the identification information of the first gNB may also be the identifier of the inference function (AIMlInferenceFunction) corresponding to the first gNB, which is not limited in this application.
[0282] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingRequest object shown in Table 4, the identification information of the target model can be represented by the value of the role-related attribute mLEntityToLoadRef of the MLEntityLoadingRequest object, and the identification information of the first gNB can be represented by the value of the attribute aIMlInferenceFunctionToLoadMlEntityRef, but this application does not make specific limitations on this.
[0283] It should be understood that in this embodiment, the MLEntityLoadingRequest includes the identification information of the first gNB. The above MLEntityLoadingRequest is used to instruct the EMS to deploy the target inference model for the first gNB, and at the same time, allow the second gNB and the third gNB within the deployment range of the target model to use the target model for inference.
[0284] Optionally, the first gNB is determined based on the resources and / or capabilities of multiple gNBs within the deployment range of the target model. The NMS can query and obtain their respective resources and capabilities from multiple gNBs, and multiple gNBs can also actively report their resources and capabilities. This application does not limit this.
[0285] In one possible implementation, the rules for the NMS to determine the target gNB may be: the first gNB has more inference resources than other gNBs, the inference speed of the first gNB is faster than other gNBs, the inference energy consumption of the first gNB is lower than that of other gNBs, the average inference result transmission delay of the first gNB within a preset time period is shorter than that of other gNBs, etc. This application does not limit this.
[0286] S802: The EMS creates an MLEntityLoadingRequest instance and an MLEntityLoadingProcess instance.
[0287] It should be understood that in this embodiment, the MLEntityLoadingProcess instance created by the EMS is a model deployment process instance for the first gNB indicated by the NMS in S601.
[0288] S803: The EMS sends a response message of MLEntityLoadingRequest to the NMS, where the response message includes the identifier of the MLEntityLoadingRequest instance and the identifier of the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the response message.
[0289] S804 . The NMS sends a deployment process status request of the target model to the EMS, including the identifier of the MLEntityLoadingRequest instance. Correspondingly, the EMS receives the request.
[0290] In another implementation of S804 , the target model deployment process status request sent by the NMS to the EMS includes an identifier of the MLEntityLoadingProcess instance.
[0291] S805: The EMS reports the target model deployment process status to the NMS, including the identifier of the MLEntityLoadingRequest instance and the progressStatus value corresponding to the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the target model deployment process status.
[0292] In another implementation of S805 , the EMS reports the deployment process status of the target model to the NMS, including the identifier of the MLEntityLoadingProcess instance and the progressStatus value corresponding to the MLEntityLoadingProcess instance.
[0293] S806: The NMS sends a control message for the target model deployment process to the EMS, including the identifier of the MLEntityLoadingRequest instance and a control instruction. Correspondingly, the EMS executes the control instruction.
[0294] In another implementation of S806 , the NMS sends a control message of the target model deployment process to the EMS, including the identifier of the MLEntityLoadingProcess instance and a control instruction.
[0295] S807. The EMS sends first information to the first gNB. The first information indicates a target model to be deployed. The first information includes identification information of the target model and a deployment range of the target model. In response, the first gNB receives the first information.
[0296] It should be understood that the identification information of the target model is used to indicate the deployment of the target model, and the deployment scope of the target model is used to indicate that any gNB belonging to the deployment scope can utilize the target model deployed on the first gNB.
[0297] S808. The first gNB deploys the target model based on the identification information of the target model.
[0298] It should be understood that the first gNB deploying the target model in S808 can be understood as the process of the first gNB activating the target model. The EMS can track the progress of the first gNB activating the target model through the MLEntityLoadingProcess instance created in the above S802, and execute S809 at the same time or after the first gNB activates the target model.
[0299] S809. The EMS sends second information to the second gNB and the third gNB. The second information indicates that the target model has been successfully deployed on the first gNB. The second information includes identification information of the target model and identification information of the first gNB. Correspondingly, the second gNB and the third gNB receive the second information.
[0300] Optionally, the first gNB may perform reasoning using the target model by using its own deployed target model.
[0301] Optionally, the second gNB and / or the third gNB may perform inference using the target model: the second gNB may perform inference using the target model deployed on the first gNB, and / or the third gNB may perform inference using the target model deployed on the first gNB, but this application is not limited thereto. For example, the process of the third gNB performing inference using the target model deployed on the first gNB may be as described in S810 to S812.
[0302] S810: The third gNB sends a first request to the first gNB. The first request is used to request an inference result of a target model. The first request includes identification information of the target model and input parameters of the target model. In response, the first gNB receives the first request.
[0303] It should be understood that the input parameters of the target model included in the first request sent by the third gNB to the first gNB may be parameters of the cell served by the third gNB, and this application does not limit this.
[0304] S811. The first gNB performs inference based on the target model and input parameters to obtain an inference result.
[0305] It should be understood that when the first gNB determines that the sending third gNB belongs to the deployment range of the target model, it inputs the input parameters of the target model included in the first request into its own deployed target model, performs inference, and obtains the inference result.
[0306] S812. The first gNB sends the inference result to the third gNB.
[0307] It should be understood that the process of the second gNB performing inference using the target model deployed on the first gNB is similar to S810 to S812 above and will not be repeated here.
[0308] It should also be understood that the above S804, S805, and S806 may be optional steps.
[0309] In the embodiment of the present application, by deploying the target model to the first gNB within the deployment range of the target model, all gNBs within the deployment range can use the target model for inference, which is beneficial to improving the flexibility of model deployment and model use, thereby reducing unnecessary resource waste.
[0310] Case 4: Scenario 1 + target entity includes some entities among multiple entities and the number of some entities is not unique
[0311] In this case, the target model deployment process can still be managed in batches or independently, specifically in the following ways:
[0312] 1) In the case of batch management, the deployment process can refer to the above-mentioned method 600, except that in S601, the MLEntityLoadingRequest sent by the NMS to the EMS also includes the identifiers of the first gNB, the second gNB, and the third gNB to indicate that the target model is deployed on the first gNB, the second gNB, and the third gNB. The deployment range of the target model may also include the fourth gNB, and the target model is not deployed on the fourth gNB. The process of the fourth gNB using the target model for inference can refer to S810 to S812 in method 800. The fourth gNB can request inference results from any of the first gNB, the second gNB, and the third gNB where the target model has been deployed. This application is not limited to this.
[0313] 2) In the case of independent management, its deployment process can refer to the above-mentioned method 700, with the difference that in S701, the MLEntityLoadingRequest sent by the NMS to the EMS also includes the identifiers of the first gNB, the second gNB, and the third gNB to indicate that the target model is deployed on the first gNB, the second gNB, and the third gNB. The deployment range of the target model may also include the fourth gNB, and the target model is not deployed on the fourth gNB. The process of the fourth gNB using the target model for inference can refer to S810 to S812 in method 800. The fourth gNB can request inference results from any of the first gNB, the second gNB, and the third gNB where the target model has been deployed. This application is not limited to this.
[0314] Case 5: Scenario 2 + target entity includes all entities in multiple entities + independent management
[0315] FIG9 is a schematic flow chart of a model deployment method 900 provided in an embodiment of the present application. The method 900 includes:
[0316] S901: The NMS sends an MLEntityLoadingPolicy to the EMS. The MLEntityLoadingPolicy includes identification information of the target model, the deployment policy of the target model, and the deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingPolicy.
[0317] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingPolicy object, the deployment policy of the target model can be represented by the value of the attribute policyForLoading, and the identification information of the target model can also be represented by the value of the attribute inferenceType or the value of the attribute mLEntityId, but this application does not limit this.
[0318] S902: EMS creates an MLEntityLoadingPolicy instance.
[0319] S903: EMS creates multiple MLEntityLoadingProcess instances.
[0320] In one possible implementation, when the EMS determines that the first gNB, the second gNB, and the third gNB all meet the deployment policy of the target model, a first MLEntityLoadingProcess instance is created for deploying the target model for the first gNB, a second MLEntityLoadingProcess instance is created for deploying the target model for the second gNB, and a third MLEntityLoadingProcess instance is created for deploying the target model for the third gNB.
[0321] Optionally, under the condition that the first gNB, the second gNB, and the third gNB simultaneously meet the deployment policy of the target model, the first MLEntityLoadingProcess instance, the second MLEntityLoadingProcess instance, and the third MLEntityLoadingProcess instance may be created simultaneously; if the first gNB, the second gNB, and the third gNB meet the deployment policy of the target model at different times, the first MLEntityLoadingProcess instance, the second MLEntityLoadingProcess instance, and the third MLEntityLoadingProcess instance may be created according to the order in which the first gNB, the second gNB, and the third gNB meet the deployment policy of the target model. This application is not limited to this.
[0322] For example, if the first MLEntityLoadingProcess instance is created before the second MLEntityLoadingProcess instance, the time when the EMS sends the first information to the first gNB may also be earlier than the time when the EMS sends the first information to the second gNB. This application is not limited to this.
[0323] In a possible implementation, the deployment process management mode of the target model may be a default mode. In this embodiment, the deployment process management mode of the target model is independent management by default.
[0324] In another possible implementation, the deployment process management mode of the target model is indicated by the NMS, and the indication information can be indicated by the infFuncBatchIndicator included in the MLEntityLoadingPolicy. In this embodiment, the infFuncBatchIndicator indicates that the deployment process management mode of the target model is independent management.
[0325] In another possible implementation, the deployment process management mode of the target model is determined by the EMS. In this embodiment, the EMS determines that the deployment process management mode of the target model is independent management.
[0326] S904: The EMS sends the identifier of the MLEntityLoadingPolicy instance and multiple identifiers corresponding to multiple MLEntityLoadingProcess instances to the NMS. Correspondingly, the NMS receives the message.
[0327] Optionally, the EMS may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S902, or may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S903. The identifier of the MLEntityLoadingPolicy instance and multiple identifiers corresponding to multiple MLEntityLoadingProcess instances may be carried in the same message or in different messages, which is not limited in this application.
[0328] S905. The NMS sends a status request for the MLEntityLoadingProcess instances corresponding to all or some of the gNBs in the multiple gNBs to the EMS, including the identifier of the MLEntityLoadingPolicy instance and the identifiers of all or some of the requested gNBs. Correspondingly, the EMS receives the request.
[0329] In another implementation of S905, the NMS sends a status request of the MLEntityLoadingProcess instances corresponding to all or part of the multiple gNBs to the EMS, including the identifiers of the MLEntityLoadingProcess instances corresponding to all or part of the requested gNBs in the multiple gNBs.
[0330] S906. The EMS reports the status of the MLEntityLoadingProcess instances corresponding to all or some of the gNBs to the NMS, including the identifier of the MLEntityLoadingPolicy instance, the identifier of all or some of the gNBs, and the progressStatus value of the MLEntityLoadingProcess instances corresponding to all or some of the gNBs. The NMS receives the status accordingly.
[0331] In another implementation of S906, the EMS reports to the NMS the status of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs, including the identifiers of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs and the progressStatus values of the MLEntityLoadingProcess instances corresponding to all or part of the target gNBs.
[0332] S907. The NMS sends a control message for the target model deployment process for all or some of the multiple target gNBs to the EMS. The message includes the identifier of the MLEntityLoadingPolicy instance, the identifiers of all or some of the gNBs to be controlled, and a control instruction for the MLEntityLoadingProcess instance corresponding to each gNB. The EMS executes the control instruction accordingly.
[0333] In another implementation of S907, the NMS sends a control message of the target model deployment process of all or part of the gNBs in the multiple gNBs to the EMS, including the identifiers of the MLEntityLoadingProcess instances corresponding to all or part of the gNBs + a control instruction for the MLEntityLoadingProcess instance corresponding to each gNB.
[0334] S908. The EMS sends first information to each of the first gNB, the second gNB, and the third gNB. The first information indicates a target deployment model and includes identification information of the target model. Correspondingly, the first gNB, the second gNB, and the third gNB receive the first information.
[0335] S909. The first gNB, the second gNB, and the third gNB deploy the target model based on the identification information of the target model.
[0336] S910: The first gNB, the second gNB, and the third gNB perform inference based on the target model.
[0337] It should be understood that S905, S906, and S907 are optional steps.
[0338] The beneficial effects of this embodiment are similar to those of the above method 700 and will not be described in detail.
[0339] Optionally, in scenario 2, when the deployment process management mode of the target model is determined by the EMS, if the EMS detects that each gNB meets the deployment policy of the target model at the same time, the target model can be deployed to multiple gNBs in a batch management manner. In this case, the EMS in S903 above can only create one MLEntityLoadingProcess instance, and in S905, S906, and S907, the interaction can be completed only through the identifier of the MLEntityLoadingPolicy instance or the identifier of the MLEntityLoadingProcess instance. After the model is successfully deployed, the EMS can continue to execute S908 to trigger subsequent processes, which will not be described in detail in this application.
[0340] Case 6: Scenario 2 + target entity includes some entities among multiple entities
[0341] FIG10 is a schematic flow chart of a model deployment method 1000 provided in an embodiment of the present application. The method 1000 includes:
[0342] S1001: The NMS sends an MLEntityLoadingPolicy to the EMS. The MLEntityLoadingPolicy includes identification information of a target model, a deployment policy of the target model, and a deployment scope of the target model. Correspondingly, the EMS receives the MLEntityLoadingPolicy.
[0343] It should be understood that when there are gNBs in the deployment range of the target model that do not require the actual deployment of the target model, the gNBs that actually require the deployment of the target model can be specified by the NMS or determined by the EMS itself.
[0344] When the gNB on which the target model is actually deployed is specified by the NMS, the MLEntityLoadingPolicy may also include the identifier of the gNB specified by the NMS on which the target model is actually deployed. The number of gNBs on which the target model is actually deployed may be one or more, and this application does not limit this.
[0345] If the target model actually needs to be deployed on multiple gNBs, S1003 to S1007 in this embodiment can be replaced by S903 to S907 described above. In this embodiment, assuming that the target model actually needs to be deployed on one gNB, the NMS specifies that the target model be deployed on the first gNB. Optionally, in S1001, the MLEntityLoadingPolicy sent by the NMS to the EMS also includes the identification information of the first gNB.
[0346] Optionally, the deployment scope of the target model can be represented by the value of the role-related attribute aIMlInferenceFunctionRef of the MLEntityLoadingPolicy object, the deployment policy of the target model can be represented by the value of the attribute policyForLoading, the identification information of the target model can be represented by the value of the attribute inferenceType or the value of the attribute mLEntityId, and the identification information of the first gNB specified by the NMS that needs to actually deploy the target model can be represented by the value of the attribute aIMlInferenceFunctionToLoadMlEntityRef, but this application is not limited to this.
[0347] S1002. EMS creates an MLEntityLoadingPolicy instance.
[0348] S1003. EMS creates an MLEntityLoadingProcess instance.
[0349] In one possible implementation, upon the EMS determining that the first gNB satisfies a deployment policy for the target model, an MLEntityLoadingProcess instance is created for deploying the target model for the first gNB.
[0350] S1004: The EMS sends the identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance to the NMS. Correspondingly, the NMS receives the message.
[0351] Optionally, the EMS may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S902, or may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S1003. The identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance may be carried in the same message or in different messages, which is not limited in this application.
[0352] S1005 . The NMS sends a deployment status request of the target model to the EMS, including the identifier of the MLEntityLoadingPolicy instance. Correspondingly, the EMS receives the request.
[0353] In another implementation of S1005 , the NMS sends a deployment status request of the target model to the EMS, including the identifier of the MLEntityLoadingProcess instance.
[0354] S1006. The EMS reports the deployment status of the target model in multiple gNBs to the NMS, including the identifier of the MLEntityLoadingPolicy instance and the progressStatus value of the MLEntityLoadingProcess instance. The NMS receives the status accordingly.
[0355] In another implementation of S1006 , the EMS reports the deployment status of the target model to the NMS, including the identifier of the MLEntityLoadingProcess instance and the progressStatus value of the MLEntityLoadingProcess instance.
[0356] S1007: The NMS sends a control message for the target model deployment process to the EMS, including the identifier of the MLEntityLoadingPolicy instance and the control instruction of the MLEntityLoadingProcess instance. Correspondingly, the EMS executes the control instruction.
[0357] In another implementation of S1007 , the NMS sends a control message of the target model deployment process to the EMS, including the identifier of the MLEntityLoadingProcess instance and the control instruction of the MLEntityLoadingProcess instance.
[0358] S1008. The EMS sends first information to each first gNB. The first information indicates a target deployment model and includes identification information of the target model. Correspondingly, the first gNB receives the first information.
[0359] S1009. The first gNB deploys the target model based on the identification information of the target model.
[0360] It should be understood that the first gNB deploying the target model in S1008 can be understood as the process of the first gNB activating the target model. The EMS can track the progress of the first gNB activating the target model through the MLEntityLoadingProcess instance created in the above S1002, and execute S1010 at the same time or after the first gNB activates the target model.
[0361] S1010. The EMS sends second information to the second gNB and the third gNB. The second information indicates that the target model has been successfully deployed on the first gNB. The second information includes identification information of the target model and identification information of the first gNB. Correspondingly, the second gNB and the third gNB receive the second information.
[0362] Optionally, the first gNB may perform reasoning using the target model by using its own deployed target model.
[0363] Optionally, the second gNB and / or the third gNB may perform inference using the target model: the second gNB may perform inference using the target model deployed on the first gNB, and / or the third gNB may perform inference using the target model deployed on the first gNB, but this application is not limited thereto. For example, the process of the third gNB performing inference using the target model deployed on the first gNB may be as described in S810 to S812.
[0364] S1011. The third gNB sends a first request to the first gNB. The first request is used to request an inference result of a target model. The first request includes identification information of the target model and input parameters of the target model. In response, the first gNB receives the first request.
[0365] It should be understood that the input parameters of the target model included in the first request sent by the third gNB to the first gNB may be parameters of the cell served by the third gNB, and this application does not limit this.
[0366] S1012. The first gNB performs inference based on the target model and input parameters to obtain an inference result.
[0367] It should be understood that when the first gNB determines that the sending third gNB belongs to the deployment range of the target model, it inputs the input parameters of the target model included in the first request into its own deployed target model, performs inference, and obtains the inference result.
[0368] S1013. The first gNB sends the inference result to the third gNB.
[0369] It should be understood that the process of the second gNB performing inference using the target model deployed on the first gNB is similar to S1010 to S1012 above and will not be repeated here.
[0370] It should be understood that S1005, S1006, and S1007 are optional steps.
[0371] The beneficial effects of this embodiment are similar to those of the above method 800 and will not be described in detail.
[0372] It should be understood that the size of the serial numbers of the various steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0373] The model deployment method of an embodiment of the present application is described in detail above in conjunction with Figures 5 to 10. The communication device of an embodiment of the present application will be described in detail below in conjunction with Figures 11 and 12.
[0374] FIG11 shows a communication device 1100 provided in an embodiment of the present application. The communication device 1100 includes a transceiver module 1101 and a processing module 1102 .
[0375] In a first possible implementation, the communication device 1100 is used to implement the steps and processes corresponding to the above-mentioned second device (eg, EMS).
[0376] Among them, the transceiver module 1101 is used to: receive the identification information of the target model and the deployment range of the target model, and the deployment range of the target model is used to indicate multiple entities that need to use the target model for reasoning; the processing module 1102 is used to: deploy the target model based on the identification information of the target model and the deployment range of the target model.
[0377] Optionally, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple reasoning functions; or identification information of multiple entities.
[0378] Optionally, the transceiver module 1101 is further used to: send first information to a target entity, where the first information is used to instruct the target entity to deploy a target model, and the target entity includes all or part of the multiple entities.
[0379] Optionally, the target entity includes all entities in the multiple entities, and the first information includes identification information of the target model.
[0380] Optionally, the target entity includes some entities among the multiple entities, and the first information includes identification information of the target model and a deployment range of the target model.
[0381] Optionally, the transceiver module 1101 is further configured to receive identification information of a target entity.
[0382] Optionally, the transceiver module 1101 is further used to: send second information, where the second information is used to indicate that the target model is deployed successfully, and the second information includes identification information of the target model and identification information of the target entity.
[0383] Optionally, the transceiver module 1101 is also used to: send first identification information and identification information of the first process, the first identification information is associated with the message carrying the identification information of the target model, there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.
[0384] Optionally, the transceiver module 1101 is further used to: receive first indication information, where the first indication information is used to indicate a deployment process management method of the target model.
[0385] Optionally, the deployment process management method is batch management.
[0386] Optionally, the deployment process management mode of the target model is batch management by default.
[0387] Optionally, the first process corresponds to an overall deployment process of the target model.
[0388] Optionally, the transceiver module 1101 is also used to: send third information, the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, the status including: running, canceling, canceled, paused or ended.
[0389] Optionally, the transceiver module 1101 is also used to: receive fourth information, the fourth information is used to control the overall deployment process of the target model, the fourth information includes first identification information and control instructions for the first process, or identification information and control instructions of the first process, the control instructions include cancel, pause or start.
[0390] Optionally, the target model's deployment process management mode is set to independent management by default.
[0391] Optionally, the deployment process management mode is independent management.
[0392] Optionally, the first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of an entity.
[0393] Optionally, the transceiver module 1101 is also used to: send fifth information, the fifth information including the first identification information, the identification information of each entity corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, the identification information of all or part of the processes, and the respective status of all or part of the processes, the status including: running, canceling, canceled, paused or ended.
[0394] Optionally, the transceiver module 1101 is also used to: receive sixth information, the sixth information is used to control all or part of the multiple processes, the sixth information includes first identification information, identification information of the entity corresponding to all or part of the processes, and control instructions for each process in all or part of the processes; or, identification information of all or part of the processes, and control instructions.
[0395] Optionally, the identification information of the target model is sent through a request message or a configuration message.
[0396] Optionally, the deployment scope of the target model is sent via a request message or a configuration message.
[0397] In a second possible implementation, the communication device 1100 is used to implement the steps and processes corresponding to the above-mentioned first device (eg, NMS).
[0398] Among them, the processing module 1102 is used to: determine the identification information of the target model and the deployment range of the target model, and the deployment range of the target model is used to indicate multiple entities that need to use the target model for reasoning; the transceiver module 1101 is used to: send the identification information of the target model and the deployment range of the target model.
[0399] Optionally, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple reasoning functions; or identification information of multiple entities.
[0400] Optionally, the transceiver module 1101 is further configured to: send identification information of a target entity, where the target entity includes some entities among the multiple entities, and the target entity is determined based on resources and / or capabilities of the multiple entities.
[0401] Optionally, the transceiver module 1101 is also used to: send first identification information and identification information of the first process, the first identification information is associated with the message carrying the identification information of the target model, there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.
[0402] Optionally, the transceiver module 1101 is further used to: receive first indication information, where the first indication information is used to indicate a deployment process management method of the target model.
[0403] Optionally, the deployment process management method is batch management.
[0404] Optionally, the deployment process management mode of the target model is batch management by default.
[0405] Optionally, the first process corresponds to an overall deployment process of the target model.
[0406] Optionally, the transceiver module 1101 is also used to: send third information, the third information includes the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, the status including: running, canceling, canceled, paused or ended.
[0407] Optionally, the transceiver module 1101 is also used to: receive fourth information, the fourth information is used to control the overall deployment process of the target model, the fourth information includes first identification information and control instructions for the first process, or identification information and control instructions of the first process, the control instructions include cancel, pause or start.
[0408] Optionally, the deployment process management mode of the target model is set to independent management by default.
[0409] Optionally, the deployment process management mode is independent management.
[0410] Optionally, the first process includes multiple processes, and any one of the multiple processes corresponds to the model deployment of an entity.
[0411] Optionally, the transceiver module 1101 is also used to: send fifth information, the fifth information including the first identification information, the identification information of each entity corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, the identification information of all or part of the processes, and the respective status of all or part of the processes, the status including: running, canceling, canceled, paused or ended.
[0412] Optionally, the transceiver module 1101 is also used to: receive sixth information, the sixth information is used to control all or part of the multiple processes, the sixth information includes first identification information, identification information of the entity corresponding to all or part of the processes, and control instructions for each process in all or part of the processes; or, identification information of the entity corresponding to all or part of the processes, and control instructions.
[0413] Optionally, the identification information of the target model is sent through a request message or a configuration message.
[0414] Optionally, the deployment scope of the target model is sent via a request message or a configuration message.
[0415] In a third possible implementation, the communication device 1100 is used to implement the steps and processes corresponding to the above-mentioned third device (e.g., entity, gNB).
[0416] Among them, the transceiver module 1101 is used to: receive the first information, the first information is used to indicate the deployment of the target model, the first information includes the deployment scope of the target model and the identification information of the target model, the deployment scope of the target model is used to indicate multiple entities that need to use the target model for reasoning; the processing module 1102 is used to: deploy the target model based on the first information.
[0417] Optionally, the deployment scope of the target model includes one or more of the following: information of a geographical area; identification information of a network area; identification information of multiple reasoning functions; or identification information of multiple entities.
[0418] Optionally, the transceiver module 1101 is also used to: receive a first request, the first request is used to request the inference result of the target model, the first request includes the identification information of the target model and the input parameters of the target model; the processing module 1102 is also used to: obtain the inference result through the target model based on the input parameters.
[0419] Optionally, the transceiver module 1101 is further used to: send the inference result.
[0420] Optionally, the transceiver module 1101 is further used to: receive second information, where the second information is used to indicate that the target model is deployed successfully, and the second information includes identification information of the target model and identification information of the target entity.
[0421] It should be understood that the device 1100 here is embodied in the form of a functional module. The term "module" here can refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combined logic circuit and / or other suitable components that support the described functions. In an optional example, those skilled in the art will understand that the device 1100 can be specifically the first device, the second device or the third device in the above-mentioned embodiment, or the functions described in the above-mentioned embodiment can be integrated in the device 1100, and the device 1100 can be used to execute the various processes and / or steps corresponding to the first device, the second device or the third device in the above-mentioned method embodiment. To avoid repetition, they are not described here.
[0422] The apparatus 1100 has the function of implementing the corresponding steps performed by the first apparatus, the second apparatus, or the third apparatus in the above method. The above functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0423] In an embodiment of the present application, the device 1100 in FIG. 11 may also be a chip or a chip system, such as a system on chip (SoC).
[0424] Figure 12 shows a schematic block diagram of a communication device 1200 provided in an embodiment of the present application. The device 1200 includes a processor 1201, a transceiver 1202, and a memory 1203. The processor 1201, the transceiver 1202, and the memory 1203 communicate with each other via an internal connection path. The memory 1203 is used to store instructions, and the processor 1201 is used to execute the instructions stored in the memory 1203 to control the transceiver 1202 to send and / or receive signals.
[0425] It should be understood that apparatus 1200 may be specifically the first apparatus, second apparatus, or third apparatus in the aforementioned embodiments, and may be configured to execute the various steps and / or processes corresponding to the first apparatus, second apparatus, or third apparatus in the aforementioned method embodiments. Optionally, the memory 1203 may include read-only memory and random access memory, and provide instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store device type information. The processor 1201 may be configured to execute instructions stored in the memory, and when the processor 1201 executes the instructions stored in the memory, the processor 1201 is configured to execute the various steps and / or processes of the aforementioned method embodiments. The transceiver 1202 may include a transmitter and a receiver. The transmitter may be configured to implement the various steps and / or processes corresponding to the aforementioned transceiver for performing a transmitting action, and the receiver may be configured to implement the various steps and / or processes corresponding to the aforementioned transceiver for performing a receiving action.
[0426] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0427] During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor executes the instructions in the memory, and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.
[0428] The present application also provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to implement the method shown in the above method embodiment.
[0429] The present application also provides a computer program product, which includes computer program code (also referred to as computer program, instruction). When the computer program runs on a computer, the computer can execute the method shown in the above method embodiment.
[0430] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0431] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0432] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0433] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0434] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0435] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program codes.< / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc>
Claims
1. A model deployment method, characterized in that: The method comprises: Receiving identification information of a target model and a deployment scope of the target model, where the deployment scope of the target model is used to indicate a plurality of entities that need to be inferred using the target model; The target model is deployed based on the identification information of the target model and the deployment scope of the target model.
2. The method according to claim 1, characterized in that The deployment scope of the target model includes one or more of the following: Geographical information; Identification information of the network area; identification information of multiple inference functions; or, Identification information for multiple entities.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: Send first identification information and identification information of a first process, wherein the first identification information is associated with a message carrying the identification information of the target model, and a mapping relationship exists between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.
4. The method according to claim 3, characterized in that The deployment process management method of the target model is batch management.
5. The method according to claim 4, characterized in that The first process corresponds to the overall deployment process of the target model.
6. The method according to claim 4 or 5, characterized in that: The method further comprises: Send third information, where the third information includes the first identification information and the state of the first process, or the identification information of the first process and the state of the first process, where the state includes: running, canceling, canceled, paused, or ended.
7. The method according to any one of claims 4 to 6, characterized in that The method further comprises: Receive fourth information, where the fourth information is used to control the overall deployment process of the target model, where the fourth information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, where the control instruction includes canceling, pausing, or starting.
8. The method according to claim 3, characterized in that The deployment process management mode of the target model is independent management.
9. The method according to claim 8, characterized in that The first process includes multiple processes, and any one of the multiple processes corresponds to a model deployment of an entity.
10. The method according to claim 8 or 9, characterized in that: The method further comprises: sending fifth information, the fifth information including the first identification information, identification information of each entity corresponding to all or part of the multiple processes, and respective states of all or part of the processes; or, The identification information of all or part of the processes, and the status of all or part of the processes, wherein the status includes: running, canceling, canceled, paused or ended.
11. The method according to any one of claims 8 to 10, characterized in that The method further comprises: receiving sixth information, the sixth information being used to control all or part of the multiple processes, the sixth information including the first identification information, identification information of entities corresponding to the all or part of the processes, and control instructions for each of the all or part of the processes; or, The identification information of all or part of the process, and the control instructions.
12. The method according to any one of claims 1 to 11, characterized in that The method further comprises: First indication information is received, where the first indication information is used to indicate a deployment process management method of the target model.
13. The method according to any one of claims 4 to 7, characterized in that The deployment process management mode of the target model is batch management by default.
14. The method according to any one of claims 8 to 11, characterized in that The deployment process management mode of the target model is independent management by default.
15. A model deployment method, characterized in that: The method comprises: Determine identification information of a target model and a deployment scope of the target model, where the deployment scope of the target model is used to indicate a plurality of entities that need to be inferred using the target model; The identification information of the target model and the deployment scope of the target model are sent.
16. The method according to claim 15, characterized in that The deployment scope of the target model includes one or more of the following: Geographical information; Identification information of the network area; identification information of multiple inference functions; or, Identification information for multiple entities.
17. The method according to claim 15 or 16, characterized in that The method further comprises: Send first identification information and identification information of a first process, wherein the first identification information is associated with a message carrying the identification information of the target model, and a mapping relationship exists between the first identification information and the first process, and the first process is used to manage the deployment process of the target model.
18. The method according to claim 17, characterized in that The deployment process management method of the target model is batch management.
19. The method according to claim 18, characterized in that The first process corresponds to the overall deployment process of the target model.
20. The method according to claim 18 or 19, characterized in that The method further comprises: Send third information, where the third information includes the first identification information and the state of the first process, or the identification information of the first process and the state of the first process, where the state includes: running, canceling, canceled, paused, or ended.
21. The method according to any one of claims 18 to 20, characterized in that The method further comprises: Receive fourth information, where the fourth information is used to control the overall deployment process of the target model, where the fourth information includes the first identification information and a control instruction for the first process, or the identification information of the first process and the control instruction, where the control instruction includes canceling, pausing, or starting.
22. The method according to claim 17, characterized in that The deployment process management mode of the target model is independent management.
23. The method according to claim 22, characterized in that The first process includes multiple processes, and any one of the multiple processes corresponds to a model deployment of an entity.
24. The method according to claim 22 or 23, characterized in that The method further comprises: sending fifth information, the fifth information including the first identification information, identification information of each entity corresponding to all or part of the multiple processes, and respective states of all or part of the processes; or, The identification information of all or part of the processes, and the status of all or part of the processes, wherein the status includes: running, canceling, canceled, paused or ended.
25. The method according to any one of claims 22 to 24, characterized in that The method further comprises: receiving sixth information, the sixth information being used to control all or part of the multiple processes, the sixth information including the first identification information, identification information of entities corresponding to the all or part of the processes, and control instructions for each of the all or part of the processes; or, The identification information of the entity corresponding to the whole or part of the process, and the control instruction.
26. The method according to any one of claims 15 to 25, characterized in that The method further comprises: First indication information is received, where the first indication information is used to indicate a deployment process management method of the target model.
27. The method according to any one of claims 18 to 21, characterized in that The deployment process management mode of the target model is batch management by default.
28. The method according to any one of claims 22 to 25, characterized in that The deployment process management mode of the target model is independent management by default.
29. The method according to any one of claims 1 to 28, characterized in that The deployment scope of the target model is sent via a request message or a configuration message.
30. A communication device, characterized in that: include: A module for executing the method according to any one of claims 1 to 14, or a module for executing the method according to any one of claims 15 to 29.
31. A communication device, characterized in that: include: A processor, wherein the processor is coupled to a memory, the memory stores computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory, so that the processor executes the method according to any one of claims 1 to 14, or executes the method according to any one of claims 15 to 29.
32. A communication system, characterized in that: The method comprises a first device and a second device, wherein the second device is used to execute the method according to any one of claims 1 to 14, and the first device is used to execute the method according to any one of claims 15 to 29.
33. A computer-readable storage medium, characterized in that: Used to store a computer program, the computer program comprising instructions for implementing the method according to any one of claims 1 to 14, or instructions for executing the method according to any one of claims 15 to 29.
34. A computer program product, comprising computer program code, characterized in that: When the computer program code runs on a computer, the computer is enabled to implement the method according to any one of claims 1 to 14 or to execute the method according to any one of claims 15 to 29.
Citation Information
Patent Citations
Model deployment method and communication device
CN120200903A
Model deployment method and device, storage medium and electronic equipment
CN112395072A
AI model deployment method and system and storage medium
CN115392332A
Model deployment method and device and electronic equipment
CN115934110A
AI model management method and device, network node and storage medium
CN116391384A