Model deployment method and communication apparatus
By receiving and utilizing the identification information of the target model group, deploying and obtaining the description files of the model group for joint training and/or joint testing, the problem of joint inference cannot be realized in the prior art is solved, and efficient model deployment and joint inference functions are achieved.
Patent Information
- Application Number
- PCT/CN2024/137847
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-09
- Publication Date
- 2025-06-26
AI Technical Summary
The prior art cannot effectively deploy multiple models that can jointly implement functions, resulting in the inability to implement joint inference functions.
By receiving the identification information of the target model group, deploying a model group that has been jointly trained and/or joint tested based on the identification information, and obtaining a description file about the model group so that the model can understand its joint reasoning method during the deployment process.
The deployment of the target model group that can jointly implement inference functions is realized, so that the deployed model can realize joint inference based on the description files obtained during the deployment process, improving the efficiency and intelligence of model deployment.
Smart Images

Figure CN2024137847_26062025_PF_FP_ABST
Abstract
Description
Model deployment method and communication device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 21, 2023, with application number 202311777250.1 and application name “Model Deployment Method and Communication Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communications, and in particular to a model deployment method and a communication device. Background Art
[0003] To improve network intelligence and automation, artificial intelligence (AI) and machine learning (ML) technologies are increasingly being applied. For example, access network equipment may include a service module (such as inference) that implements functions based on ML models. These modules can then implement corresponding functions based on ML models, thereby increasing the intelligence of these devices.
[0004] Generally speaking, each model can independently support a specific type of function. Some model deployment methods can deploy one or more independent models for devices that need to deploy models (such as the access network devices mentioned above). However, in practice, the implementation of a function may require the cooperation of multiple models. For example, the output of one model is used as the input of another model to form a sequence of interconnected models, or the output is provided in parallel by multiple models, or the parallel output of multiple models is combined and used as the input of another model, and the function is jointly implemented by such multiple models. These multiple models can be models for joint training and / or joint testing.
[0005] However, currently, there is no deployment method for such models that can jointly realize functions. Summary of the Invention
[0006] The present application provides a model deployment method and a communication device, which are conducive to deploying a model group that can realize joint reasoning functions.
[0007] In the first aspect, the present application provides a model deployment method, which can be executed by a second device or a component in the second device (such as a chip, a chip system, etc.), and the second device has the ability to provide model loading services. The second device can be a network management system, a network element management system, a network device or a wireless access network device, or any other device or system that can provide model loading services. This application does not specifically limit this. The method includes: receiving identification information of a target model group, the target model group including multiple models, and the multiple models are jointly trained and / or jointly tested; based on the identification information of the target model group, deploying the target model group.
[0008] It should be understood that the purpose of deploying a model is to make the deployed model available on the entity that implements the reasoning function. In one possible implementation, when deploying a model, the entity that implements the reasoning function will read the model's description file, which may be a description file for the model's capabilities and needs to be obtained based on the model's identification information. Although the existing model deployment method can deploy multiple models, on the one hand, multiple models may be trained or tested independently, and the purpose of joint reasoning cannot be achieved. On the other hand, even if these multiple models are jointly trained and tested, since the multiple models are deployed in an independent manner based on their respective model identifiers, only the description files of each model can be read, and the description file of the model group cannot be obtained, that is, the joint method between multiple models cannot be obtained, that is, multiple models can each implement the reasoning functions they support, but cannot implement the joint reasoning function.
[0009] In an embodiment of the present application, in the process of deploying the target model group, the identification information of the target model group indicates that the target model group deployed is a target model group that has been jointly trained and / or jointly tested and can jointly implement the reasoning function, so that a description file about the target model group can be obtained based on the identification information of the target model group during the deployment process. The description file may include the cooperation method of multiple models in the target model group when jointly implementing the reasoning function. For example, the cooperation method can be to use the output of one model in the multiple models as the input of another model to form a mutually linked model sequence, or the multiple models provide outputs in parallel, or the parallel outputs of some models in the multiple models are merged and used as the input of another part of the model, etc. In this way, the multiple models of the deployed target model group can realize joint reasoning based on the description file obtained during the deployment process. It can be seen that the model deployment method provided by the present application can realize the deployment of the target model group that can jointly implement the reasoning function, which is conducive to the target model group to realize joint reasoning.
[0010] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: sending first identification information and identification information of a first process, the first identification information being associated with a message carrying the identification information of the target model group, the first identification information being mapped to the first process, and the first process being used to manage the deployment process of the target model group.
[0011] In a possible implementation, the identification information of the target model group is carried by a request message, and the first identification information may be identification information of a request instance created based on the request message.
[0012] In another possible implementation, the identification information of the target model group is carried by a configuration message, and the first identification information may be identification information of an instance created based on the configuration message.
[0013] In combination with the first aspect, in some implementations of the first aspect, first indication information is received, where the first indication information is used to indicate a deployment process management method for the target model group.
[0014] In a possible implementation, the first indication information indicates that the deployment process management mode of the target model group is batch management.
[0015] In another possible implementation, the deployment process management mode of the target model group is batch management by default, that is, no explicit instruction is required from either party, for example, it may be agreed upon in a protocol.
[0016] It should be understood that the target model group includes multiple models, and "batch management" can be understood as deploying these multiple models as a whole during the model deployment process.
[0017] In combination with the first aspect, in certain implementations of the first aspect, when the deployment process management mode of the target model group is batch management, the first process corresponds to the overall deployment process of the target model group.
[0018] Optionally, the number of process instances corresponding to the first process may be one, but this application does not limit this.
[0019] In the embodiment of the present application, the entire deployment process of the target model group can be completed through one process, which is beneficial to saving resources and energy consumption of the second device.
[0020] In combination with the first aspect, in some implementations of the first aspect, the method further includes: sending first information, the first information including the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, the status including: running, canceling, canceled, paused or ended.
[0021] In a possible implementation manner, the first information is actively reported by the second device.
[0022] In another possible implementation, the first information is reported by the second device in response to a request from the first device. The method further includes: before sending the first information, receiving a request message for requesting the status of the first process, the request message including the first identification information and the status of the first process, or the identification information of the first process and the status of the first process.
[0023] In an embodiment of the present application, when the model deployment process management method is batch management, the overall deployment process of the target model group can be uniquely identified by the first identification information or the identification information of the first process. When the second device reports the status of the first process, it can indicate the correspondence between the currently reported status information and the model deployment process based on the first identification information or the identification information of the first process. When a target model group is deployed and the target model group includes multiple models, the second device may also only include the status information of one process when reporting the first information, which is beneficial to reducing the signaling overhead between the second device and the device interacting with the second device.
[0024] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: receiving second information, the second information being used to control the overall deployment process of the target model group, the second information including the first identification information and control instructions for the first process, or the identification information of the first process and the control instructions, the control instructions including canceling, pausing or starting.
[0025] In an embodiment of the present application, when the model deployment process management mode is batch management, the overall deployment process of the target model group can be controlled by a control instruction, which is conducive to improving the model deployment efficiency.
[0026] In combination with the first aspect, in certain implementations of the first aspect, the deployment process management mode of the target model group is independent management by default.
[0027] In combination with the first aspect, in some implementations of the first aspect, the first indication information indicates that the deployment process management mode of the target model group is independent management.
[0028] "Independent management" can be understood as creating an independent process for deployment process management for each model included in the target model group.
[0029] In combination with the first aspect, in certain implementations of the first aspect, the first process includes multiple processes, the multiple processes correspond one-to-one to the identification information of each model in the multiple models, the multiple processes respectively correspond to the deployment processes of the multiple models, and the identification information of the first process includes multiple identification information corresponding to the multiple processes.
[0030] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: sending third information, the third information including the first identification information, identification information of each model corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, identification information of all or part of the processes, and the respective status of all or part of the processes, the status including: running, canceling, canceled, paused or ended.
[0031] In a possible implementation, the third information is actively reported by the second device.
[0032] In another possible implementation, the third information is reported by the second device in response to a request from the first device. The method further includes: before sending the third information, receiving a message for requesting the status of all or some of the multiple processes, the message including the first identification information, identification information of each model corresponding to all or some of the multiple processes, or identification information of all or some of the processes.
[0033] In combination with the first aspect, in some implementations of the first aspect, the method further includes: receiving fourth information, the fourth information being used to control the deployment process of all or part of the multiple models, the fourth information including the first identification information, identification information of each model in the all or part of the models, and control instructions for the deployment process of each model in the all or part of the models; or, identification information of the deployment process of each model in the all or part of the models, and the control instructions.
[0034] In an embodiment of the present application, through an independently managed model deployment method, the first process may include multiple processes, the multiple processes correspond one-to-one to the identification information of each model in the multiple models, and the multiple processes correspond to the deployment processes of the multiple models respectively. In this way, the first device can respectively obtain and control the deployment status of each model in the target model group, and can reflect the different deployment status of each model. In the case that some models in the target model group fail to be deployed, it is possible to determine which models have failed to be deployed, and based on the identification information of some models that failed to be deployed, these models can be redeployed without redeploying the entire target model group, which is conducive to improving management accuracy and reducing unnecessary waste of resources.
[0035] In combination with the first aspect, in some implementations of the first aspect, the identification information of the target model group is sent via a request message or a configuration message.
[0036] When the identification information of the target model group is sent through a request message, the request message can be used to indicate the deployment of the target model group; when the identification information of the target model group is sent through a configuration message, the configuration message can also include the deployment strategy of the target model group.
[0037] Optionally, the deployment strategy of the target model group can be a requirement for the reasoning capability of the target model group, or it can be the parameters of each functional entity involved in the model deployment process, etc. The deployment strategy can be determined by the first device, or it can be reported by the entity that assumes the reasoning function. This application does not limit this. Exemplarily, the deployment strategy can be that the accuracy of the target model group is greater than 90%, or that the preset time point or preset period for deploying the target model group is reached in time, or that a certain network performance indicator is monitored to be lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations on this.
[0038] In combination with the first aspect, in some implementations of the first aspect, the method further includes: receiving identification information of each model in the target model group.
[0039] In a possible implementation, the identification information of each model in the target model group is sent through a request message or a configuration message.
[0040] In another possible implementation, each functional entity participating in the deployment of the target model group stores a mapping relationship between the identification information of the target model group and the identification information of each model in the target model group. In this way, after any party obtains the identification information of the target model group, it can search locally based on the identification information of the target model group, or request the identification information of each model in the target model group from other entities.
[0041] On the second aspect, the present application further provides a model deployment method, which can be executed by a first device or a component in the first device (such as a chip, a chip system, etc.), the first device has the ability to call a model loading service, the first device can be a network management system, a network element management system, a network device or a wireless access network device, or any other device or system that has the ability to call a model loading service, and the present application does not make specific limitations on this. The method includes: determining identification information of a target model group, the target model group includes multiple models, and the multiple models are jointly trained and / or jointly tested; sending identification information of the target model group.
[0042] Optionally, the way in which the first device determines the identification information of the target model group may include: the first device determines the identification information of the target model group corresponding to the demand based on the demand reported by the target reasoning function entity; or, the first device determines the identification information of the target model group corresponding to the target reasoning function based on the operation status of the monitored target reasoning function entity, etc. This application does not limit this.
[0043] Optionally, the identification information of the target model group may be stored locally in the first device, or may be obtained by the first device from other devices, and this application does not limit this.
[0044] In combination with the second aspect, in certain implementations of the second aspect, the method further includes: receiving first identification information and identification information of a first process, the first identification information being associated with a message carrying the identification information of the target model group, the first identification information being mapped to the first process, and the first process being used to manage the deployment process of the target model group.
[0045] In combination with the second aspect, in some implementations of the second aspect, the method further includes: sending first indication information, where the first indication information is used to indicate a deployment process management method for the target model group.
[0046] In combination with the second aspect, in some implementations of the second aspect, the deployment process management method is batch management.
[0047] In combination with the second aspect, in certain implementations of the second aspect, the deployment process management mode of the target model group is batch management by default.
[0048] In combination with the second aspect, in some implementations of the second aspect, the first process corresponds to the overall deployment process of the target model group.
[0049] In combination with the second aspect, in some implementations of the second aspect, the method also includes: receiving first information, the first information including the first identification information and the status of the first process, or the identification information of the first process and the status of the first process, the status including: running, canceling, canceled, paused or ended.
[0050] In combination with the second aspect, in certain implementations of the second aspect, the method further includes: sending second information, the second information being used to control the overall deployment process of the target model group, the second information including the first identification information and control instructions for the first process, or the identification information of the first process and the control instructions, the control instructions including canceling, pausing or starting.
[0051] In combination with the second aspect, in certain implementations of the second aspect, the deployment process management mode of the target model group is independent management by default.
[0052] In combination with the second aspect, in some implementations of the second aspect, the deployment process management method is independent management.
[0053] In combination with the second aspect, in certain implementations of the second aspect, the first process includes multiple processes, the multiple processes correspond one-to-one to the identification information of each model in the multiple models, the multiple processes respectively correspond to the deployment processes of the multiple models, and the identification information of the first process includes multiple identification information corresponding to the multiple processes.
[0054] In combination with the second aspect, in some implementations of the second aspect, the method further includes: receiving third information, the third information including the first identification information, identification information of each model corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, identification information of all or part of the processes, and the respective status of all or part of the processes, the status including: running, canceling, canceled, paused or ended.
[0055] In combination with the second aspect, in some implementations of the second aspect, the method further includes: sending fourth information, wherein the fourth information is used to control the deployment process of all or part of the multiple models, and the fourth information includes the first identification information, the identification information of each model in the all or part of the models, and the control instructions for the deployment process of each model in the all or part of the models; or, the identification information of the deployment process of each model in the all or part of the models, and the control instructions.
[0056] In combination with the second aspect, in some implementations of the second aspect, the identification information of the target model group is sent via a request message or a configuration message.
[0057] In combination with the second aspect, in some implementations of the second aspect, the method further includes: sending identification information of each model in the target model group.
[0058] In a third aspect, the present application provides a communication device comprising a module for executing the method as described in any one of the first aspect or the second aspect.
[0059] In a fourth aspect, the present application further provides a communication device, comprising a processor coupled to a memory and configured to execute instructions in the memory to implement the method of any possible implementation of the first or second aspect described above. Optionally, the device further comprises a memory. Optionally, the device further comprises a communication interface, the processor coupled to the communication interface.
[0060] In a fifth aspect, a processor is provided, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is configured to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method of any possible implementation of the first or second aspect.
[0061] In a specific implementation, the processor may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be a transistor, a gate circuit, a trigger, or various logic circuits. The input signal received by the input circuit may be, for example, but not limited to, received and input by a receiver, and the signal output by the output circuit may be, for example, but not limited to, output to and transmitted by a transmitter. The input circuit and the output circuit may be the same circuit, which functions as an input circuit and an output circuit at different times. The embodiments of the present application do not limit the specific implementation of the processor and various circuits.
[0062] In a sixth aspect, a processing device is provided, comprising a processor and a memory. The processor is configured to read instructions stored in the memory and receive signals via a receiver and transmit signals via a transmitter to execute the method of any possible implementation of the first or second aspect.
[0063] Optionally, there are one or more processors and one or more memories.
[0064] Optionally, the memory may be integrated with the processor, or the memory may be provided separately from the processor.
[0065] In the specific implementation process, the memory can be a non-transitory memory, such as a read-only memory (ROM), which can be integrated with the processor on the same chip or can be set on different chips. The embodiments of the present application do not limit the type of memory and the setting method of the memory and the processor.
[0066] It should be understood that related data interaction processes, such as sending indication information, can be the process of outputting indication information from the processor, and receiving capability information can be the process of receiving input capability information from the processor. Specifically, the output data of the processing can be output to the transmitter, and the input data received by the processor can come from the receiver. The transmitter and receiver can be collectively referred to as a transceiver.
[0067] The processing device in the sixth aspect mentioned above can be a chip. The processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading the software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0068] In the seventh aspect, a computer program product is provided, which includes: a computer program (also referred to as code, or instructions), which, when executed, enables a computer to execute a method in any possible implementation of the first or second aspect.
[0069] In an eighth aspect, a computer-readable storage medium is provided, which stores a computer program (also referred to as code, or instructions) which, when run on a computer, enables the computer to execute a method in any possible implementation of the first or second aspect above.
[0070] The beneficial effects and possible implementation methods of the third to eighth aspects can be referred to the description of the first and second aspects and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] FIG1 is a model workflow provided by an embodiment of the present application;
[0072] FIG2 is a working diagram of a network resource model provided in an embodiment of the present application;
[0073] FIG3 is a working diagram of a network resource model provided by an embodiment of the application;
[0074] FIG4 is another working diagram of a network resource model provided in an embodiment of the application;
[0075] FIG5 is a schematic diagram of a communication system applicable to the embodiment of the application;
[0076] FIG6 is a schematic flow chart of a model deployment method provided in an embodiment of the present application;
[0077] FIG7 is a schematic flow chart of another model deployment method provided in an embodiment of the present application;
[0078] FIG8 is a schematic flow chart of another model deployment method provided in an embodiment of the present application;
[0079] FIG9 is a schematic flow chart of another model deployment method provided in an embodiment of the present application;
[0080] FIG10 is a schematic flow chart of a model deployment method provided in an embodiment of the present application;
[0081] FIG11 is a schematic block diagram of a communication device provided in an embodiment of the present application;
[0082] FIG12 is a schematic block diagram of another communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0083] The technical solution in this application will be described below with reference to the accompanying drawings.
[0084] Before introducing this application, the following points are explained.
[0085] First, in the embodiments described below, various terms and abbreviations, such as baseline data and differential data, are provided for ease of description and should not be construed as limiting this application. This application does not exclude the possibility of defining other terms in existing or future protocols that can achieve the same or similar functions.
[0086] Second, the first, second and various numerical numbers in the embodiments shown below are only used for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
[0087] Third, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c, where a, b, c can be single or multiple.
[0088] In order to improve the intelligence and automation level of the network, artificial intelligence (AI) and machine learning (ML) technologies are gradually being applied to promote network intelligence.
[0089] Figure 1 exemplifies the AI / ML workflow, which mainly includes the training phase, simulation phase, deployment phase and inference phase. It should be understood that, in general, the AI / ML workflow is implemented in sequence according to the training phase, simulation phase, deployment phase and inference phase, but in some possible implementations, the AI / ML model may execute the inference phase after the training phase and / or simulation phase, or skip the simulation phase after the training phase to execute the deployment phase and inference phase, or enter the training phase after the simulation phase or inference phase. Among them, the deployment phase is the process of loading the model into the inference function. In one possible implementation, after the model has gone through the training phase and the simulation phase, it is loaded into the corresponding inference function through the deployment phase to enter the inference phase to execute the inference process.
[0090] In some examples, the role that provides the ML entity loading management service (MnS) is called a producer, and the role that calls the ML entity loading management service is called a consumer. Two sub-use cases and solutions for these two sub-use cases are defined for the deployment phase.
[0091] The two sub-use cases include:
[0092] Sub-use case 1: Consumer requested ML entity loading.
[0093] Sub-use case 2: The consumer configures the policy for the producer, and the producer triggers the model loading (control of producer-initiated ML entity loading).
[0094] Solutions for two sub-use cases:
[0095] The current solution defines three information object classes (IOCs): MLEntityLoadingRequest, MLEntityLoadingPolicy, and MLEntityLoadingProcess, which represents the ML model loading process. MLEntityLoadingRequest and MLEntityLoadingProcess are used in the scenario described in sub-use case 1, while MLEntityLoadingPolicy and MLEntityLoadingProcess are used in the scenario described in sub-use case 2.
[0096] The following describes in detail the attributes of the above three information object classes.
[0097] 1.MLEntityLoadingRequest< <ioc>>
[0098] MLEntityLoadingRequest< <ioc>> represents an ML model loading request created by a consumer. The consumer uses this IOC to request the producer to load the ML model into the target inference function. The properties of this IOC are shown in Table 1.
[0099] Table 1
[0100] The requestStatus attribute indicates the status of the request and can be any of the following: NOT_STARTED, LOADING_IN_PROGRES, SUSPENDED, FINISHED_SUCCESS, FINISHED_FAILED, or CANCELLED. The cancelRequest attribute indicates whether the consumer canceled the ML entity loading request and can be TRUE or FALSE. The mLEntityToLoadRef role attribute indicates the identifier of the model to be loaded in the ML model loading request created by the consumer.
[0101] The values for "Support Qualifier," "Readability," and "Writability" in the table above can be understood as follows: "Support Qualifier" can take the following values: mandatory (M), optional (O), and conditionally mandatory (CM). A value of M indicates that the attribute is mandatory, a value of O indicates that the attribute is optional, and a value of CM indicates that the attribute is conditionally mandatory, meaning that the attribute is mandatory only when certain conditions are met. "Readability" can take the following values: true (T) and false (F). A value of T indicates that the attribute is readable, while a value of F indicates that the attribute is not readable. "Writability" can also take the following values: true (T) and false (F). A value of T indicates that the attribute is writable, while a value of F indicates that the attribute is not writable. This explanation also applies to the following tables and will not be repeated here.
[0102] 2.MLEntityLoadingPolicy< <ioc>>
[0103] MLEntityLoadingPolicy< <ioc>> represents the ML model loading policy set by the consumer for the producer. The consumer uses this IOC to set the conditions for triggering ML model loading for the producer. The producer triggers model loading only when the ML model meets the conditions. The properties of this IOC are shown in Table 2.
[0104] Table 2
[0105] The inferenceType attribute indicates the inference type of the initial model associated with the ML model loading policy set by the consumer for the producer. This attribute is conditionally required, meaning that it is required only when the model associated with the ML model loading policy set by the consumer for the producer is the initial training model. It should be understood that the initial training model is the model being trained for the first time and does not yet have a model identifier. Therefore, the inference type of the initial training model can be used to identify the model.
[0106] The mLEntityId attribute is used to indicate the identifier of the retrained model associated with the ML model loading policy set by the consumer for the producer. This attribute is conditionally required. That is, when the model associated with the ML model loading policy set by the consumer for the producer is a retrained model, this attribute is required. It should be understood that an mLEntityId uniquely identifies a model that has already been initially trained.
[0107] The policyForLoading attribute is used to indicate the ML model loading policy set by the consumer to the producer. The policy can be a series of thresholds. For example, the policy can be that the accuracy of the ML model is greater than 90%, or a preset time point or preset period is reached, or a certain network performance indicator is lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule).
[0108] 3. MLEntityLoadingProcess <ioc>>
[0109] MLEntityLoadingProcess< <ioc>> represents an ML model loading process. In the scenario described in sub-use case 1 above, the producer can use this IOC to instantiate one or more ML model loading processes for each ML model loading request from the consumer. In the scenario described in sub-use case 2 above, the producer can use this IOC to instantiate one or more ML model loading processes and associate these one or more ML model loading processes with the ML model loading policy set by the consumer for the producer. The properties of this IOC are shown in Table 3.
[0110] Table 3
[0111] The progressStatus attribute indicates the status of the ML model loading process. The value of this attribute can be any of the following: RUNNING, CANCELLING, SUSPENDED, FINISHED, or CANCELLED. The cancelProcess attribute indicates whether the consumer cancels the ML entity loading process. The value is true (TRUE) or false (FALSE). The suspendProcess attribute indicates whether the consumer suspends the ML entity loading process. The value is true (TRUE) or false (FALSE). The resumeProcess attribute indicates whether the consumer resumes the ML entity loading process. The value is true (TRUE) or false (FALSE).
[0112] The role-related attribute MLEntityLoadingRequestRef is the loading request identifier associated with the ML entity loading process and is required when the ML entity loading process corresponds to the scenario described in Sub-Use Case 1. MLEntityLoadingPolicyRef is the loading policy identifier associated with the ML entity loading process and is required when the ML entity loading process corresponds to the scenario described in Sub-Use Case 2. LoadedMLEntityRef is the model identifier associated with the ML entity loading process.
[0113] Figure 2 illustrates a diagram of the network resource model (NRM) relationships for the IOCs corresponding to Tables 1, 2, and 3. As shown in Figure 2, AiMlInferenceFunction represents a function that can execute ML inference and is also an IOC. In the figure, the lines connecting the MLEntityLoadingRequest, MLEntityLoadingProcess, and MLEntityLoadingPolicy objects with the AiMlInferenceFunction object are solid diamonds on the AiMlInferenceFunction object side. These solid diamonds represent inclusion relationships, meaning that the AiMlInferenceFunction object includes the MLEntityLoadingRequest, MLEntityLoadingProcess, and MLEntityLoadingPolicy objects. Furthermore, the lines connecting the MLEntityLoadingRequest object, the MLEntityLoadingProcess object, and the MLEntity object are arrows on the MLEntity object side. These arrows represent the association relationship, meaning that the MLEntityLoadingRequest object and the MLEntityLoadingProcess object are each associated with the MLEntity object. Therefore, this can be further understood as the AiMlInferenceFunction object containing the MLEntity object. In the figure, the lines connecting the AiMlInferenceFunction object and the MLEntity object are solid diamonds on the AiMlInferenceFunction object side.
[0114] In addition, the numbers and symbols on the lines connecting the various objects in the diagram represent the cardinality relationships between the intent object attributes that have inclusion and association relationships. For example, in the line between the MLEntityLoadingRequest object and the AiMlInferenceFunction object shown in Figure 2, the cardinality on the MLEntityLoadingRequest object side is "*", and the cardinality on the AiMlInferenceFunction object side is "1". On the line between the MLEntityLoadingRequest object and the MLEntity object, the cardinality on both the MLEntityLoadingRequest object and the MLEntity object is "1". This means that an AiMlInferenceFunction object can contain multiple MLEntityLoadingRequest objects, while each MLEntityLoadingRequest object can only be associated with one MLEntity object. The other lines in Figure 2 and the connection rules of the subsequent NRM diagrams can also be interpreted according to this rule and will not be repeated here.
[0115] That is, when loading a model for an inference function, only one model can be loaded through a single model loading request initiated by the consumer in the above sub-use case 1. Generally speaking, each model can independently support a specific type of function, but in actual applications, the implementation of an inference function may require the cooperation of multiple models. For example, the output of one model can be used as the input of another model to form an interconnected model sequence, or the output of multiple models can be provided in parallel, or the parallel output of multiple models can be combined and used as the input of another model to jointly implement the function. These multiple models can be jointly trained and / or jointly tested.
[0116] According to the design of Figure 2, although multiple models can be loaded through multiple model loading requests, on the one hand, the models loaded by multiple model loading requests may be independently trained or tested, and the purpose of joint reasoning cannot be achieved. On the other hand, even if these multiple models are jointly trained and tested, since the multiple models are deployed in an independent manner based on their respective model identifiers, only the description files of each model can be read, and the joint method between the multiple models cannot be obtained. In other words, multiple models can each implement the reasoning functions they support, but cannot implement the joint reasoning function.
[0117] In view of this, the present application provides a model deployment method and communication device, which can load a set of target model groups for the target reasoning function, and the target model group includes multiple models for joint training and / or joint testing. The description file about the target model group can be read based on the identification information of the target model group during the deployment process. The description file may include the cooperation method of the multiple models in the target model group when jointly implementing the reasoning function. In this way, the multiple models of the deployed target model group can implement joint reasoning based on the description file obtained during the deployment process, thereby achieving the purpose of deploying a target model group that can support joint reasoning for the target reasoning function.
[0118] Below, the design of the IOC proposed in this application is first described in detail. There are two design methods as follows.
[0119] Design 1:
[0120] 1. Add an object that can model a joint inference model group.
[0121] For example, the object may be MLEntityCoordinationGroup< <ioc>>, but this application does not limit this. <ioc>> can represent a target model group, which may include multiple models that can be jointly trained and / or jointly tested to achieve joint reasoning.
[0122] Optionally, the model can be an AI model or an ML model (also referred to as an ML entity), which is not limited in this application. In the subsequent description of this application, the model is only used as an example of an ML model, but this example does not constitute a limitation of this application.
[0123] It is worth noting that one MLEntityCoordinationGroup is associated with at least two ML models.
[0124] In one possible implementation, MLEntityCoordinationGroup< <ioc>The properties of > are shown in Table 4.
[0125] Table 4
[0126] Optionally, the role-related attribute memberMLEntityRefList in Table 4 above may represent a list of model identifiers in a target model group, but this application does not specifically limit this, and the model identifiers in the target model group may also be embodied in other forms.
[0127] 2. Add the following two properties to the model loading request object created by the consumer:
[0128] Attribute 1: used to indicate the management method during the model deployment process (such as batch management or independent management);
[0129] Attribute 2: Used to indicate that the model loading request created by the consumer is used to request the deployment of a target model group, where multiple models in the target model group can implement joint reasoning.
[0130] For example, a model loading request created by a consumer can be generated through MLEntityLoadingRequest< <ioc>> indicates that Table 5 exemplarily shows an MLEntityLoadingRequest< provided in an embodiment of the present application <ioc>>Contained attributes.
[0131] Table 5
[0132] In one possible implementation, the above-mentioned attribute 1 may be a groupBatchIndicator, which may be an optional attribute. The groupBatchIndicator may be used to indicate the management method during the model deployment process, for example, batch management of the target model group or independent management of multiple models in the target model group. The above-mentioned attribute 2 may be a role-related attribute mLEntityCoordinationGroupToLoadRef, which represents the identifier of the target model group that the ML model loading request created by the consumer needs to load through mLEntityCoordinationGroupToLoadRef. At the same time, the support qualifier of mLEntityToLoadRef is modified to CM. In this way, when the model that needs to be loaded in the ML model loading request created by the consumer is not the target model group, mLEntityToLoadRef is a required attribute. When the model that needs to be loaded in the ML model loading request created by the consumer is the target model group, mLEntityCoordinationGroupToLoadRef is a required attribute. However, it should be understood that this application does not specifically limit the specific forms of attributes 1 and 2.
[0133] 3. Add the following two properties to the model loading strategy object set by the consumer to the producer:
[0134] Attribute 3: used to indicate the management method during the model deployment process (such as batch management or independent management);
[0135] Attribute 4: Used to indicate that the model associated with the model loading strategy set by the consumer to the producer is the target model group. Multiple models in the target model group can achieve joint reasoning.
[0136] For example, the model loading policy set by the consumer to the producer can be set by MLEntityLoadingPolicy< <ioc>> indicates that Table 6 exemplarily shows an MLEntityLoadingPolicy provided in an embodiment of the present application. <ioc>>Contained attributes.
[0137] Table 6
[0138] In one possible implementation, the above-mentioned attribute 3 may be groupBatchIndicator, which may be an optional attribute and may indicate the management method during the model deployment process; the above-mentioned attribute 4 may be a role-related attribute mLEntityCoordinationGroupToLoadRef, which represents the identifier of the target model group associated with the ML model loading strategy set by the consumer for the producer through mLEntityCoordinationGroupToLoadRef. This attribute may be a conditionally required attribute, that is, when the model associated with the ML model loading strategy set by the consumer for the producer is the target model group, mLEntityCoordinationGroupToLoadRef is a required attribute, but it should be understood that this application does not specifically limit the specific forms of attributes 3 and 4.
[0139] Optionally, the value restrictions for the identifier of the retrained model corresponding to the model loading policy set by the consumer for the producer can also be modified. In one possible implementation, the identifier of the retrained model corresponding to the model loading policy set by the consumer for the producer can be represented by mLEntityId. As shown in Figure 2, the value of mLEntityId is restricted to 1, and an MLEntityLoadingPolicy object can only be associated with one MLEntity object. In one possible implementation of this design, the value restrictions on mLEntityId are removed, allowing an MLEntityLoadingPolicy object to be associated with one or more MLEntity objects.
[0140] 4. Add the following two properties to the model loading process object:
[0141] Attribute 5: used to associate the deployment process of the target model group;
[0142] Attribute 6: used to associate the deployment process of the target model group and the deployment process of multiple models in the target model group.
[0143] Optionally, the model loading process may also be referred to as the model deployment process, which is not limited in this application.
[0144] For example, the model loading process object can be accessed through MLEntityLoadingProcess< <ioc>> indicates that Table 7 exemplarily shows an MLEntityLoadingProcess< provided in an embodiment of the present application <ioc>>Contained attributes.
[0145] Table 7
[0146] In one possible implementation, attribute 5 may be LoadedMLEntityCoordinationGroupRef, which is a conditionally required attribute. For example, if the model to be loaded corresponding to sub-use case 1 or sub-use case 2 is the target model group, LoadedMLEntityCoordinationGroupRef is a required attribute. Attribute 6 may be LoadedMemberMLEntityRef, which is also a conditionally required attribute. For example, if the model to be loaded corresponding to sub-use case 1 or sub-use case 2 is the target model group, and the management method during model deployment is to independently manage multiple models in the target model group, this attribute is a required attribute. However, it should be understood that this application does not limit the specific forms of attributes 5 and 6.
[0147] Figure 3 illustrates an exemplary diagram of the network resource model relationships of various IOCs corresponding to Design 1 provided in an embodiment of the present application. Compared to Figure 2, Figure 3 adds an MLEntityCoordinationGroup object and defines the cardinality relationships between the MLEntityLoadingRequest object, the MLEntityLoadingProcess object, and the MLEntityLoadingPolicy object, respectively, and the MLEntityCoordinationGroup object.
[0148] In the arrow line from the MLEntityLoadingRequest object to the MLEntityCoordinationGroup object and the MLEntity object, the cardinality on the side close to the MLEntityLoadingRequest object is "1", and the cardinality on the side close to the MLEntityCoordinationGroup object and the MLEntity object is "0..1", indicating that one MLEntityLoadingRequest object can be associated with one MLEntityCoordinationGroup object, or one MLEntityLoadingRequest object can be associated with one MLEntity object.
[0149] On the arrow line from the MLEntityLoadingPolicy object to the MLEntityCoordinationGroup object and the MLEntity object, the cardinality is "1" on the side close to the MLEntityLoadingPolicy object, and the cardinality is "0..1" on the side close to the MLEntityCoordinationGroup object and the MLEntity object, indicating that an MLEntityLoadingPolicy object can be associated with an MLEntityCoordinationGroup object, or an MLEntityLoadingPolicy object can be associated with an MLEntity object.
[0150] On the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityCoordinationGroup object and the MLEntity object, the cardinality on the side closest to the MLEntityCoordinationGroup and MLEntity objects is "0..1", indicating that the MLEntityLoadingProcess object can be associated with one MLEntityCoordinationGroup object, or one MLEntity object. On the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityCoordinationGroup object, the cardinality on the side closest to the MLEntityLoadingProcess object is "1..*", indicating that when the MLEntityLoadingProcess object is associated with an MLEntityCoordinationGroup object, there can be at least one MLEntityLoadingProcess object ("at least one" can also be described as "one or more"). On the arrow line pointing from the MLEntityLoadingProcess object to the MLEntity object, the cardinality on the side closest to the MLEntityLoadingProcess object is "1", indicating that when the MLEntityLoadingProcess object is associated with an MLEntity object, there can be only one MLEntityLoadingProcess object.
[0151] In addition, on the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityLoadingPolicy object, and on the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityLoadingRequest object, the cardinality close to the MLEntityLoadingProcess object has been changed from "*" (shown in Figure 2) to "1..*", indicating that at least one MLEntityLoadingProcess object can be associated with one MLEntityLoadingPolicy object, and at least one MLEntityLoadingProcess object can be associated with one MLEntityLoadingRequest object.
[0152] In one possible implementation, based on the above-mentioned design 1, when deploying a target model group for an inference function, it can be implemented by associating an MLEntityLoadingRequest object with an MLEntityCoordinationGroup object, or it can be implemented by associating an MLEntityLoadingPolicy object with an MLEntityCoordinationGroup object.
[0153] Design 2:
[0154] 1. Modification of the model loading request object created by the consumer:
[0155] 1) Modify the value of the identifier of the model to be loaded in the model loading request created by the consumer;
[0156] 2) New attribute 7: used to indicate that the model load request created by the consumer is for requesting the deployment of a target model group, where multiple models in the target model group can achieve joint reasoning;
[0157] 3) New attribute 8: used to indicate the management method during the model deployment process (such as batch management or independent management).
[0158] For example, a model loading request created by a consumer can be generated through MLEntityLoadingRequest< <ioc>> indicates that Table 8 exemplarily shows another MLEntityLoadingRequest< provided in an embodiment of the present application <ioc>>Contained attributes.
[0159] Table 8
[0160] In one possible implementation, the identifier of the model that needs to be loaded by the model loading request created by the consumer can be represented by mLEntityToLoadRef. As shown in Figure 2, the value of mLEntityToLoadRef is limited to 1, and one MLEntityLoadingRequest object can only be associated with one MLEntity object. In a possible implementation of this design, the value restriction on mLEntityToLoadRef is cancelled, that is, one MLEntityLoadingRequest object can be associated with one or more MLEntity objects. In addition, a new attribute 7 is added to indicate that the mLEntityToLoadRef corresponding to one or more MLEntity objects associated with an MLEntityLoadingRequest object belongs to a target model group that can realize the joint reasoning function. In this way, when the model that needs to be loaded by the ML model loading request created by the consumer is the target model group, the above attribute 7 is a required attribute. Optionally, the above attribute 7 can be represented by mLEntityGroupInfo, but this application does not limit this.
[0161] In one possible implementation, the above-mentioned attribute 8 may be a groupBatchIndicator, which may be an optional attribute. The groupBatchIndicator may be used to indicate the management method during the model deployment process. For example, the management method may be batch management of the target model group, or independent management of multiple models in the target model group, but this application does not make any specific limitations on this.
[0162] 2. Modification of the model loading strategy object set by the consumer to the producer:
[0163] 1) Modify the value of the identifier of the retraining model corresponding to the model loading strategy set by the consumer to the producer;
[0164] 2) New attribute 9: used to indicate that the model load request created by the consumer is for requesting the deployment of a target model group, where multiple models in the target model group can achieve joint reasoning;
[0165] 3) New attribute 10: used to indicate the management method during the model deployment process (such as batch management or independent management).
[0166] For example, the model loading policy set by the consumer to the producer can be set by MLEntityLoadingPolicy< <ioc>> indicates that Table 9 exemplarily shows another MLEntityLoadingPolicy provided in an embodiment of the present application. <ioc>>Contained attributes.
[0167] Table 9
[0168] In one possible implementation, the identifier of the retrained model corresponding to the model loading policy set by the consumer for the producer can be represented by mLEntityId. As shown in Figure 2, the value of mLEntityId is limited to 1, and one MLEntityLoadingPolicy object can only be associated with one MLEntity object. In a possible implementation of this design, the value restriction on mLEntityId is removed, so that one MLEntityLoadingPolicy object can be associated with one or more MLEntity objects. In addition, a new attribute 9 is added to indicate that the mLEntityId corresponding to one or more MLEntity objects associated with an MLEntityLoadingPolicy object belongs to a target model group that can realize the joint reasoning function. In this way, when the model associated with the ML model loading policy set by the consumer for the producer is the target model group, the above attribute 9 is a required attribute. Optionally, the above attribute 9 can be represented by mLEntityGroupInfo, but this application does not limit this.
[0169] In one possible implementation, the above-mentioned attribute 10 may be a groupBatchIndicator, which may be an optional attribute. The groupBatchIndicator may be used to indicate a management method during the model deployment process. For example, the management method may be batch management of the target model group, or independent management of multiple models in the target model group, but this application does not make any specific limitations on this.
[0170] 3. Modification of the model loading process object: Modify the value of the model identifier associated with the loading process.
[0171] For example, the model loading process object can be accessed through MLEntityLoadingProcess< <ioc>> indicates that Table 10 exemplarily shows an MLEntityLoadingProcess< provided in an embodiment of the present application <ioc>>Contained attributes.
[0172] Table 10
[0173] In one possible implementation, the model identifier associated with the loading process can be represented by LoadedMLEntityRef. In some examples, the value of LoadedMLEntityRef is limited to 1, that is, one MLEntityLoadingProcess object corresponds to only one LoadedMLEntityRef. In a possible implementation of this design, the value restriction on LoadedMLEntityRef is removed, so that one MLEntityLoadingProcess object can correspond to one or more LoadedMLEntityRef. For example, when the deployment process management mode is batch management, one MLEntityLoadingProcess object can correspond to multiple LoadedMLEntityRefs, and the multiple LoadedMLEntityRefs in this case respectively correspond to the identification information of multiple models in the target model group; when the deployment process management mode is independent management, one MLEntityLoadingProcess object can correspond to one LoadedMLEntityRef, and the LoadedMLEntityRef in this case corresponds to the identification information of a model in the target model group, but this application does not limit this. Optionally. The model in the target model group can also be called a sub-model of the target model group, which is not limited in this application.
[0174] Figure 4 illustrates a network resource model relationship diagram for various IOCs corresponding to Design 2 provided in an embodiment of the present application. As shown in Figure 4 , the arrow line from the MLEntityLoadingRequest object pointing to the MLEntity object has a cardinality of "1" on the side closest to the MLEntityLoadingRequest object and a cardinality of "1..*" on the side closest to the MLEntity object, indicating that one MLEntityLoadingRequest object can be associated with at least one MLEntity object. The arrow line from the MLEntityLoadingPolicy object pointing to the MLEntity object has a cardinality of "1" on the side closest to the MLEntityLoadingPolicy object and a cardinality of "1..*" on the side closest to the MLEntity object, indicating that one MLEntityLoadingPolicy object can be associated with at least one MLEntity object. The arrow line from the MLEntityLoadingProcess object pointing to the MLEntity object has a cardinality of "1" on the side closest to the MLEntityLoadingProcess object and a cardinality of "1..*" on the side closest to the MLEntity object, indicating that one MLEntityLoadingProcess object can be associated with at least one MLEntity object.
[0175] In addition, on the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityLoadingPolicy object, and on the arrow line pointing from the MLEntityLoadingProcess object to the MLEntityLoadingRequest object, the cardinality close to the MLEntityLoadingProcess object has been changed from "*" (shown in Figure 2) to "1..*", indicating that at least one MLEntityLoadingProcess object can be associated with one MLEntityLoadingPolicy object, and at least one MLEntityLoadingProcess object can be associated with one MLEntityLoadingRequest object.
[0176] In one possible implementation, based on the second design, when deploying a target model group for an inference function, this can be achieved by associating at least one MLEntity object with an MLEntityLoadingRequest object, or by associating at least one MLEntity object with an MLEntityLoadingPolicy object. When deploying the target model group, the number of the at least one MLEntity object can be greater than or equal to two.
[0177] Figure 5 is a schematic diagram of a communication system 500 applicable to an embodiment of the present application. As shown in Figure 5, the communication system 500 includes a first device 510, a second device 520, and a third device 530. The first device 510, the second device 520, and the third device 530 can establish communication with each other.
[0178] In one possible implementation, the role of the first device 510 can be a model loading consumer (ML loading consumer) that can call the model loading service, the role of the second device 520 can be a model loading producer (ML loading producer) that can provide model loading services, and the third device 530 can be a role that implements the inference function (ML inference) based on the model (it can also be understood as the bearer of the inference function, the executor of the inference function, and the obtainer of the inference result. In some examples, this role can be directly called the inference function), but this application does not make any specific limitations on this.
[0179] It should be understood that in the communication system shown in FIG5 , the number of the first device, the second device, or the third device can be one or more, and this application does not limit this.
[0180] In one possible scenario, the first device 510 may call the model loading service provided by the second device 520 to instruct the target model to be loaded into the designated third device 530 , and the third device 530 is responsible for inputting data into the model to obtain corresponding output.
[0181] In one possible implementation, the first device 510 (or the second device 520, or the third device 530) may include one or more of the following: a module that can implement the corresponding function of model loading consumers, a module that can implement the corresponding function of model loading producers, or a module that can implement reasoning functions based on the model. This application does not make specific limitations on this.
[0182] Optionally, the first device 510 can be a network management system (NMS), an element management system (EMS), or a radio access network (RAN) device, or any other device or system that can be used as a model to load consumers, and this application does not make specific limitations on this.
[0183] Optionally, the second device 520 can be a network management system, a network element management system, or a wireless access network device, or any other device or system that can serve as a model loading producer, and this application does not make any specific limitations on this.
[0184] Optionally, the third device 530 can be a network management system, a network element management system, or a wireless access network device, or any other device or system that can implement reasoning functions based on a model, and this application does not make any specific limitations on this.
[0185] Optionally, the above-mentioned network element management system can manage the network and be responsible for the operation, management and functional maintenance of the network. It can also be called a cross-domain management system, which is not limited in this application.
[0186] Optionally, the network element management system can be used to manage one or more network elements of a certain category, and can also be called a domain management system or a single domain management system, which is not limited in this application.
[0187] Optionally, the above-mentioned access network device can be a transmission reception point (TRP), an evolved NodeB (eNB or eNodeB) in a long term evolution (LTE) system, a home base station (for example, home evolved NodeB, or home Node B, HNB), a base band unit (BBU), or a wireless controller in a cloud radio access network (CRAN) scenario, or the access network device can be a relay station, an access point, a vehicle-mounted device, a wearable device, and an access network device in a 5G network or an access network device in a future evolved public land mobile communication network (PLMN) network, etc., it can be an access point (AP) in a wireless local area network (WLAN), it can be a gNB in a new radio (NR) system, it can be a satellite base station in a satellite communication system, etc., and the embodiments of the present application are not limited.
[0188] Optionally, the access network device may include a centralized unit (CU) node, a distributed unit (DU) node, or an access network device including a CU node and a DU node, or an access network device including a control plane CU node (CU-CP node) and a user plane CU node (CU-UP node) and a DU node. The access network device including the CU node and the DU node may split the protocol layer of the access network device, with the functions of some protocol layers being centrally controlled by the CU, and the functions of the remaining part or all of the protocol layers being distributed in the DU, which is centrally controlled by the CU. As an implementation method, the CU deployment protocol stack includes a radio resource control (RRC) layer, a packet data convergence protocol (PDCP) layer, and a service data adaptation protocol (SDAP) layer. The DU deployment protocol stack includes a radio link control (RLC) layer, a media access control (MAC) layer, and a physical layer (PHY) layer. Thus, the CU has the processing capabilities of RRC, PDCP, and SDAP. DU has the processing capabilities of RLC, MAC and PHY. The above-mentioned functional division is only an example and does not constitute a limitation on CU and DU. That is to say, there may be other ways of functional division between CU and DU, which will not be described in detail in the embodiments of the present application. The functions of CU can be implemented by one entity or by different entities. For example, the functions of CU can be further divided, for example, the control plane (CP) and the user plane (UP) are separated, that is, the control plane of CU (CU-CP) and the user plane of CU (CU-UP). For example, CU-CP and CU-UP can be implemented by different functional entities, and CU-CP and CU-UP can be coupled with DU to jointly complete the functions of access network equipment. In one possible way, CU-CP is responsible for control plane functions, mainly including RRC and PDCP-C, where PDCP-C is mainly responsible for encryption and decryption, integrity protection, data transmission, etc. of control plane data. CU-UP is responsible for user plane functions, mainly including SDAP and PDCP-U, where SDAP is mainly responsible for processing data of core network devices and mapping data flows to bearers. The PDCP-U is responsible for data plane encryption and decryption, integrity protection, header compression, sequence number maintenance, and data transmission. The CU-CP and CU-UP are connected via the E1 interface. The CU-CP represents access network equipment and connects to core network equipment via the interface between the core network and access network equipment.The DU is connected via F1-C (control plane). The CU-UP is connected to the DU via F1-U (user plane). In addition, another possible implementation is that the PDCP-C is also in the CU-UP, which is not limited in this application.
[0189] In order to make the purpose and technical solution of this application clearer and more intuitive, the model deployment method and communication device provided by the embodiment of this application will be described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0190] Below, a model deployment method provided in an embodiment of the present application is first described in detail with reference to Figures 6 to 10.
[0191] The model deployment method provided in the embodiments of this application can be implemented based on the two aforementioned designs for IOC provided in this application. This method can be applied to the communication system 500 shown in Figure 5 or other communication systems, and this application does not limit this. For example, in the embodiments of this application, the model deployment method provided in this application is described from the perspective of inter-device interaction, taking the role of the first device as a consumer, the role of the second device as a producer, and the role of the third device as a model-based reasoning function as an example.
[0192] FIG6 is a schematic flow chart of a model deployment method 600 provided in an embodiment of the present application. The method 600 includes the following steps:
[0193] S601. The first device determines identification information of a target model group, where the target model group includes multiple models, and the multiple models are jointly trained and / or jointly tested.
[0194] It should be noted that the model in this application can be a machine learning model (ML model) or can be embodied in the form of an entity, such as a machine learning entity (ML entity), and this application does not limit this.
[0195] S602: The first device sends identification information of the target model group. Correspondingly, the second device receives the identification information of the target model group.
[0196] S603: The second device deploys the target model group based on the identification information of the target model group.
[0197] It should be noted that the second device deploying the target model group can be understood as the second device loading the target model group to one or more target inference functions (for example, the third device or gNB mentioned below), where the target model group will be used to perform inference.
[0198] In an embodiment of the present application, during the process of deploying the target model group, the second device indicates through the identification information of the target model group that the target model group deployed is a target model group that has been jointly trained and / or jointly tested and can jointly implement the reasoning function, so that a description file about the target model group can be obtained based on the identification information of the target model group during the deployment process. The description file may include the coordination method of multiple models in the target model group when jointly implementing the reasoning function. For example, the coordination method may be to use the output of one model in the multiple models as the input of another model to form a mutually linked model sequence, or the multiple models provide outputs in parallel, or the parallel outputs of some models in the multiple models are merged and used as the input of another part of the model, etc. In this way, the multiple models of the deployed target model group can realize joint reasoning based on the description file obtained during the deployment process. It can be seen that the model deployment method provided by the present application can realize the deployment of the target model group that can jointly implement the reasoning function, which is conducive to the target model group to realize joint reasoning.
[0199] In one possible scenario, the purpose of the first device determining the identification information of the target model group in the above S601 is to deploy the target model group for the third device to implement the target reasoning function on the third device. Exemplarily, the way in which the first device determines the identification information of the target model group may include: the first device determines the identification information of the target model group corresponding to the target reasoning function required by the third device based on the requirements reported by the third device; or, the first device determines the identification information of the target model group corresponding to the target reasoning function required by the third device based on the monitored operation status of the third device, etc., which is not limited in this application. Optionally, the identification information of the target model group may be stored locally by the first device, or may be obtained by the first device from other devices, which is not limited in this application.
[0200] It should be understood that the purpose of deploying the model is to enable the deployed model to be available on the entity of the reasoning function (such as a third device). A possible implementation of the above S603 may be: the second device deploys the target model group for one or more third devices based on the identification information of the target model group. In some examples, the second device sends a second indication message to the third device, and the indication message is used to instruct the third device to activate the existing target model group. The indication message may include the identification information of the target model group; in other examples, the second device sends a third indication message to the third device, and the indication message is used to instruct the third device to download and activate the target model group. The indication message may include the identification information of the target model group and the storage address of the target model group.
[0201] As an optional embodiment, after S602 above, method 600 further includes: the second device sending first identification information and identification information of the first process, the first identification information being associated with a message carrying identification information of the target model group, the first identification information being mapped to the first process, and the first process being used to manage the deployment process of the target model group. Correspondingly, the first device receiving the first identification information and the identification information of the first process.
[0202] Optionally, the first identification information and the identification information of the first process may be sent via a response message, and the response message may be a response to a message carrying the identification information of the target model group, but this application does not limit this.
[0203] It should be understood that the embodiments of the present application are applicable to the scenarios corresponding to the two sub-use cases defined in the model deployment stage mentioned above. In different scenarios, the identification information of the target model group can be carried by different types of messages, so the meaning of the first identification information is different in different scenarios.
[0204] Scenario 1: Scenario corresponding to the above-mentioned sub-use case 1. In a possible implementation of this scenario, the identification information of the target model group in S601 is sent by the first device via a request message, and the request message is used to request the second device to deploy the target model group for the third device.
[0205] It should be understood that in this scenario, the first device determines the timing of sending the above-mentioned request message to the second device based on the deployment strategy of the target model group. The deployment strategy can be the requirement for the reasoning ability of the target model group, or the parameters of the various functional entities involved in the model deployment process, etc. The deployment strategy can be determined by the first device or reported by the third device. This application does not limit this. For example, the deployment strategy can be that the accuracy of the target model group is greater than 90%, or the time point or preset period for deploying the target model group is reached in time, or a certain network performance indicator of the third device is monitored to be lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations on this.
[0206] In one possible implementation, when a first device determines that a target model group and / or a third device satisfy conditions corresponding to a deployment policy, the first device sends a request message to a second device, requesting deployment of the target model group. The request message includes identification information of the target model group. In response, the second device receives the request message and creates an instance of the request message and a first process instance.
[0207] In some possible implementations, the request message further includes identification information of each model in the target model group.
[0208] It should be understood that the first process is used to manage the deployment process of the target model group. The first process can be referred to as the target model group deployment process or the target model group loading process. This application does not specifically limit the specific name of the first process. In addition, unless otherwise specified, the terms "loading" and "deployment" described in the embodiments of this application can be understood to have the same meaning and will not be repeated hereafter.
[0209] In this scenario, the IOC and attributes corresponding to the request message can have the following two examples:
[0210] In the first example, the request message may correspond to the MLEntityLoadingRequest object described in Design 1 above. The MLEntityLoadingRequest object is associated with the MLEntityCoordinationGroup object to indicate that the request message requests the deployment of a model group. The identification information of the target model group can be represented by the value of the role-related attribute mLEntityCoordinationGroupToLoadRef of the MLEntityLoadingRequest object.
[0211] Optionally, the identification information of each model in the target model group may be represented by the value of the role-related attribute memberMLEntityRefList of the MLEntityCoordinationGroup object associated with the MLEntityLoadingRequest object.
[0212] Correspondingly, the first process can be a process created based on the MLEntityLoadingProcess object described in Design 1. The request message can be identified by the value of the role-related attribute MLEntityLoadingRequestRef of the MLEntityLoadingProcess object, and the identification information of the target model group corresponding to the first process can be represented by the value of the role-related attribute LoadedMLEntityCoordinationGroupRef of the MLEntityLoadingProcess object.
[0213] In the second example, the request message may correspond to the MLEntityLoadingRequest object described in the second design above. The mLEntityGroupInfo property of the MLEntityLoadingRequest object is used to indicate that the request message requests the deployment of the target model group. The value of mLEntityGroupInfo may be used to indicate the identification information of the target model group.
[0214] Optionally, the identification information of each model in the target model group may be represented by multiple values of the role-related attribute mLEntityToLoadRef in the MLEntityLoadingRequest object.
[0215] Correspondingly, the first process can be a process created based on the MLEntityLoadingProcess object described in Design 2. The request message can be identified by the value of the role-related attribute MLEntityLoadingRequestRef of the MLEntityLoadingProcess object, and the identification information of the target model group or the identification information of multiple models in the target model group can be indicated by one or more values of the attribute LoadedMLEntityRef of the MLEntityLoadingProcess object.
[0216] In the two examples corresponding to scenario 1, the instance of the request message may be an MLEntityLoadingRequest instance, the first identification information may be the identifier of the MLEntityLoadingRequest instance, the instance of the first process may be an MLEntityLoadingProcess instance, and the identification information of the first process may be the identifier of the MLEntityLoadingProcess instance.
[0217] Scenario 2: Scenario corresponding to the above-mentioned sub-use case 2. In a possible implementation of this scenario, the identification information of the target model group in S601 is sent by the first device via a configuration message, and the configuration message also includes a deployment strategy of the target model group.
[0218] In one possible implementation, the first device sends the identification information of the target model group and the deployment policy of the target model group to the second device through a configuration message. After the second device receives the identification information of the target model group and the deployment policy of the target model group, it creates an instance of the configuration message, and after determining that the target model group and / or the third device have met the conditions corresponding to the deployment policy, it creates an instance of the first process.
[0219] In a possible implementation, the configuration message further includes identification information of each model in the target model group.
[0220] Optionally, the first device may send the identification information of the target model group and the deployment strategy of the target model group through one configuration message, or may send the identification information of the target model group and the deployment strategy of the target model group respectively through two configuration messages. This application does not limit the number, content, and sequence of configuration messages.
[0221] Optionally, the configuration strategy of the target model group in this scenario may be that the accuracy of the target model group is greater than 90%, or that the preset time point or preset period for deploying the target model group is reached, or that a certain network performance indicator of the third device is monitored to be lower than a preset threshold (such as: energy efficiency is less than 500 bits per joule), etc. This application does not make specific limitations on this.
[0222] In this scenario, the IOC and attributes corresponding to the configuration message can have the following two examples:
[0223] In the first example, the configuration message may correspond to the MLEntityLoadingPolicy object described in the first design above. The MLEntityLoadingPolicy object may be associated with the MLEntityCoordinationGroup object to indicate that the deployment policy included in the configuration message corresponds to the deployment of the target model group. The identification information of the target model group may be represented by the value of the role-related attribute mLEntityCoordinationGroupToLoadRef of the MLEntityLoadingPolicy object. The deployment policy of the target model group may be identified by the attribute policyForLoading.
[0224] In a possible implementation, the value of the role-related attribute memberMLEntityRefList of the MLEntityCoordinationGroup object associated with the MLEntityLoadingPolicy object represents the identification information of each model in the target model group.
[0225] In another possible implementation, multiple values of the mLEntityId attribute of the MLEntityLoadingPolicy object correspond to identification information of multiple models in the target model group.
[0226] Correspondingly, the first process can be a process created based on the MLEntityLoadingProcess object described in Design 1. The configuration message can be identified by the value of the role-related attribute MLEntityLoadingPolicyRef of the MLEntityLoadingProcess object, and the identification information of the target model group can be represented by the value of the role-related attribute LoadedMLEntityCoordinationGroupRef of the MLEntityLoadingProcess object.
[0227] In the second example, the configuration message may correspond to the MLEntityLoadingPolicy object described in the second design above. The identification information of the target model group may be represented by the value of the mLEntityGroupInfo attribute of the MLEntityLoadingPolicy object. The deployment policy of the target model group may be identified by the attribute policyForLoading.
[0228] Optionally, multiple values of the mLEntityId attribute may be used to indicate identification information of multiple models in the target model group.
[0229] Correspondingly, the first process is a process created based on the MLEntityLoadingProcess object described in Design 2. The configuration message can be identified by the value of the role-related attribute MLEntityLoadingPolicyRef of the MLEntityLoadingProcess object, and the identification information of multiple models in the target model group can be corresponded by multiple values of the attribute LoadedMLEntityRef of the MLEntityLoadingProcess object.
[0230] In the two examples corresponding to scenario 2, the instance of the configuration message may be an MLEntityLoadingPolicy instance, the first identification information may be the identifier of the MLEntityLoadingPolicy instance, the instance of the first process may be an MLEntityLoadingProcess instance, and the identification information of the first process may be the identifier of the MLEntityLoadingProcess instance.
[0231] In both scenarios, the first process is used to manage the deployment process of the target model group. Optionally, the deployment process management mode for the target model group can include "batch management" and "independent management," but this application does not specifically limit this. The meaning of the first process varies under different deployment process management modes. The following uses "batch management" and "independent management" as examples for detailed description.
[0232] The first method: batch management
[0233] The target model group includes multiple models. "Batch management" can be understood as treating these multiple models as a whole during the model deployment process. This means that all trained / tested models in the model group are simultaneously deployed / loaded into the target inference function. In this case, the first process corresponds to the overall deployment process for the target model group. For example, the second device creates an MLEntityLoadingProcess instance to manage the deployment process for the target model group.
[0234] It should be understood that the MLEntityLoadingProcess instance has a status attribute, progressStatus, which can take any of the following values: RUNNING, CANCELLING, SUSPENDED, FINISHED, or CANCELLED. When the progressStatus value of MLentityLoadingProcess is RUNNING, it can be understood that the model deployment process is in progress. Therefore, RUNNING can also be understood as "deploying" or "loading", which is not limited in this application.
[0235] Illustratively, the second device may send the second indication information or the third indication information to the third device while creating the MLEntityLoadingProcess instance, or after creating the MLEntityLoadingProcess instance and when the progressStatus value of the MLEntityLoadingProcess instance is RUNNING.
[0236] Optionally, the second device may report the status of the first process to the first device.
[0237] When the deployment process management mode is batch management, the number of the first process may be one, and the identification information of the first process identifies a process instance that can represent the entire deployment process of the target model group.
[0238] In one possible implementation, the second device may proactively report the status of the first process to the first device. The method further includes: the second device sending first information, the first information including first identification information and the status of the first process, or the identification information of the first process and the status of the first process. Correspondingly, the first device receives the first information.
[0239] It should be understood that the first information is used to report the status of the first process, and the first identification information is associated with the message carrying the identification information of the target model group and has a mapping relationship with the first process. Therefore, the second device can use the first identification information or the identification information of the first process to identify the first process. The status of the first process can include: running, canceling, canceled, paused, or ended.
[0240] In another possible implementation, the second device may report the status of the first process to the first device in response to a first request from the first device. Specifically, before the second device sends the first information, the second device receives a first request from the first device, where the first request is used to request the status of the first process, and the first request includes first identification information, or identification information of the first process.
[0241] Optionally, the first device may control the first process. In one possible implementation, the method further includes: the first device sending second information, the second information being used to control the overall deployment process of the target model group; and the third information including the first identification information and a control instruction for the first process, or the identification information and control instruction for the first process, the control instruction including canceling, pausing, or starting. Correspondingly, the second device receives the second information.
[0242] In an embodiment of the present application, through a batch management model deployment method, the number of first processes can be one, and the second device can maintain only one MLEntityLoadingProcess instance to achieve the deployment of the target model group. In the status request and reporting process of the first process, the target model group can be treated as a whole, and the target model group can also be controlled as a whole during the control process of the first process. There is no need to carry information about each model in the target model group, which is beneficial to reducing the signaling overhead between each device and improving the deployment efficiency of the target model group.
[0243] The second method: independent management
[0244] "Independent management" can be understood as creating an independent process for each model included in the target model group to manage the deployment process. In some examples, the multiple models can be deployed sequentially during the deployment process of the target model group, that is, the multiple models in the trained target model group are deployed sequentially to the target inference function; in other examples, the multiple models can be deployed simultaneously during the deployment process of the target model group, that is, the multiple models in the trained target model group are deployed simultaneously to the target inference function. This application does not specifically limit the deployment order of the multiple models in the target model group.
[0245] In this case, the first process includes multiple processes, each corresponding to a deployment process for multiple models in the target model group. The identification information of the first process includes multiple identification information corresponding to the multiple processes. Exemplarily, the second device creates multiple MLEntityLoadingProcess instances to manage the deployment process for the target model group. These multiple MLEntityLoadingProcess instances correspond to the multiple models in the target model group.
[0246] When the deployment process management method is independent management, the second device reports the status of the first process to the first device by identifying the deployment process of this model through the first identification information and the identification information of a model in the target model group, or the deployment process of this process can be identified by the identification information of a certain process.
[0247] In one possible implementation, the second device actively sends third information, and the third information is used to report the status of all or part of the multiple processes. The third information includes first identification information, identification information of each model corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, identification information of all or part of the processes, and the respective status of all or part of the processes.
[0248] In another possible implementation, the second device may report the status of all or some of the multiple processes to the first device in response to a second request from the first device. Specifically, before the second device sends the third information, the second device receives a second request from the first device, where the second request is used to request the status of all or some of the multiple processes, and the second request includes the first identification information, identification information of each model corresponding to all or some of the multiple processes, or identification information of all or some of the processes.
[0249] Optionally, the first device may control multiple processes included in the first process based on the control instruction. In one possible implementation, the method further includes: the first device sending fourth information, the fourth information being used to control the deployment process of all or some of the multiple models, the fourth information including the first identification information, identification information of each model in all or some of the models, and the control instruction for the deployment process of each model in all or some of the models; or identification information of the deployment process of each model in all or some of the models and the control instruction.
[0250] In an embodiment of the present application, through an independently managed model deployment method, the first process may include multiple processes, the multiple processes correspond one-to-one to the identification information of each model in the multiple models, and the multiple processes correspond to the deployment processes of the multiple models respectively. In this way, the first device can respectively obtain and control the deployment status of each model in the target model group, and can reflect the different deployment status of each model. In the case that some models in the target model group fail to be deployed, it is possible to determine which models have failed to be deployed, and based on the identification information of some models that failed to be deployed, these models can be redeployed without redeploying the entire target model group, which is conducive to improving management accuracy and reducing unnecessary waste of resources.
[0251] In one possible implementation, the deployment process management mode of the target model group is a default mode, which includes batch management or independent management, that is, no explicit instructions are required from either party. Optionally, the deployment process management mode can be agreed upon in the protocol, which is not limited in this application.
[0252] In another possible implementation, the deployment process management mode of the target model group is indicated by the first device. The implementation method may be: the first device sends first indication information to the second device, where the first indication information is used to indicate the deployment process management mode of the target model group.
[0253] Optionally, the first indication information can be represented by the value of the attribute groupBatchIndicator. In the above scenario 1, the first indication information can be represented by the value of the attribute groupBatchIndicator of the MLEntityLoadingRequest object; in the above scenario 2, the first indication information can be represented by the value of the attribute groupBatchIndicator of the MLEntityLoadingPolicy object, but this application does not limit this.
[0254] Optionally, the value of the first indication information groupBatchIndicator may include 0 and 1. When the value is 0, it indicates that the deployment process management mode of the target model group is batch management. When the value is 1, it indicates that the deployment process management mode of the target model group is independent management. This application does not specifically limit the value of groupBatchIndicator and the meaning of the value.
[0255] Below, taking the first device in Figure 6 as the NMS, the second device as the EMS, and the third device that undertakes the model inference function as the gNB as an example, for the above scenarios 1 and 2, as well as the two deployment process management methods of batch management and independent management, the following four cases are described in detail in combination with Figures 7 to 10.
[0256] Case 1: Scenario 1 + Batch Management
[0257] FIG7 is a schematic flow chart of a model deployment method 700 provided in an embodiment of the present application. The method 700 includes:
[0258] S701: When the NMS determines that the deployment policy of the target model group is satisfied, it sends an MLEntityLoadingRequest to the EMS, where the MLEntityLoadingRequest includes identification information of the target model group. Correspondingly, the EMS receives the MLEntityLoadingRequest.
[0259] Optionally, the MLEntityLoadingRequest also includes identification information of each model in the target model group.
[0260] In one possible implementation, corresponding to Design 1 above, the target model group's identification information can be indicated by the MLEntityCoordinationGroup associated with the MLEntityLoadingRequest. Alternatively, the identification information of each model in the target model group can be indicated by the value of the memberMLEntityRefList role-related attribute of the MLEntityCoordinationGroup object.
[0261] In another possible implementation, corresponding to the second design, the target model group's identification information can be indicated by the mLEntityGroupInfo attribute of the MLEntityLoadingRequest object. Alternatively, the identification information of each model in the target model group can be represented by multiple values of the mLEntityToLoadRef role-related attribute in the MLEntityLoadingRequest object.
[0262] S702: The EMS creates an MLEntityLoadingRequest instance and an MLEntityLoadingProcess instance.
[0263] It should be understood that the MLEntityLoadingProcess instance created in S702 corresponds to the overall deployment process of the target model group, that is, multiple models in the target model group are deployed and managed as a whole.
[0264] It should also be understood that the deployment process management mode of the target model group can be default or NMS-indicated. In the case of NMS-indicated, the indication information can be indicated by the groupBatchIndicator included in the MLEntityLoadingRequest. In this embodiment, the groupBatchIndicator indicates that the deployment process management mode of the target model group is batch management.
[0265] S703: The EMS sends a response message of the MLEntityLoadingRequest to the NMS, where the response message includes the identifier of the MLEntityLoadingRequest instance and the identifier of the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the response message.
[0266] S704 : The NMS sends a deployment process status request of the target model group to the EMS, including the identifier of the MLEntityLoadingRequest instance. Correspondingly, the EMS receives the request.
[0267] In another implementation of S704 , the NMS sends a deployment process status request of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance.
[0268] S705: The EMS reports the deployment process status of the target model group to the NMS, including the identifier of the MLEntityLoadingRequest instance and the progressStatus value corresponding to the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the deployment process status of the target model group.
[0269] In another implementation of S705 , the EMS reports the deployment process status of the target model group to the NMS, including the identifier of the MLEntityLoadingProcess instance and the progressStatus value corresponding to the MLEntityLoadingProcess instance.
[0270] S706: The NMS sends a control message for the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingRequest instance and a control instruction. Correspondingly, the EMS executes the control instruction.
[0271] In another implementation of S706 , the NMS sends a control message of the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance and a control instruction.
[0272] S707. The EMS sends an indication message to the gNB, indicating activation of the target model group. The indication message includes identification information of the target model group. In response, the gNB receives the indication message.
[0273] It should be understood that S707 can be understood as the process of the EMS performing model deployment and loading the target model group into the target gNB. Optionally, the number of target gNBs can be one or more, which is not limited in this application.
[0274] Optionally, the indication message may be the second indication message or the third indication message mentioned above, which is not limited in this application.
[0275] S708. The gNB activates the target model group based on the identification information of the target model group.
[0276] In one possible implementation, after S708, method 700 further includes: the gNB performing an inference function based on the target model group.
[0277] It should be understood that the above S704, S705, and S706 are all optional steps.
[0278] In this embodiment, the MLEntityLoadingRequest includes the target model group's identification information, indicating that the model deployment is targeting a specific model group. During the deployment process, the gNB can retrieve the target model group's description file based on the target model group's identification information, enabling joint inference across multiple models in the activated target model group. Furthermore, the EMS maintains an MLEntityLoadingProcess instance for this model group, tracking the gNB's activation process for the target model group, thereby saving energy and improving model deployment efficiency.
[0279] Case 2: Scenario 1 + Independent Management
[0280] FIG8 is a schematic flow chart of a model deployment method 800 provided in an embodiment of the present application. The method 800 includes:
[0281] S801: When the NMS determines that the deployment policy of the target model group is satisfied, it sends an MLEntityLoadingRequest to the EMS, where the MLEntityLoadingRequest includes identification information of the target model group. Correspondingly, the EMS receives the MLEntityLoadingRequest.
[0282] Optionally, the MLEntityLoadingRequest also includes identification information of each model in the target model group.
[0283] It should be understood that the representation of the identification information of the target model group and the identification information of each model in the target model group in method 800 may be the same as that of the above method 700, and will not be repeated here.
[0284] S802: The EMS creates an MLEntityLoadingRequest instance and multiple MLEntityLoadingProcess instances.
[0285] It should be understood that the multiple MLEntityLoadingProcess instances correspond one-to-one to the multiple models in the target model group.
[0286] It should also be understood that the deployment process management mode of the target model group can be default or NMS-indicated. In the case of NMS-indicated, the indication information can be indicated by the groupBatchIndicator included in the MLEntityLoadingRequest. In this embodiment, the groupBatchIndicator indicates that the deployment process management mode of the target model group is independent management.
[0287] S803: The EMS sends a response message of MLEntityLoadingRequest to the NMS, where the response message includes the identifier of the MLEntityLoadingRequest instance and multiple identifiers corresponding to multiple MLEntityLoadingProcess instances. Correspondingly, the NMS receives the response message.
[0288] S804 . The NMS sends a deployment process status request for all or part of the models in the target model group to the EMS, including the identifier of the MLEntityLoadingRequest instance and the identification information of all or part of the requested models. Correspondingly, the EMS receives the request.
[0289] In another implementation of S804 , the NMS sends a deployment process status request of the target model group to the EMS, including all or part of the identification information of the MLEntityLoadingProcess instance.
[0290] S805: The EMS reports the deployment process status of all or some models in the target model group to the NMS, including the identifier of the MLEntityLoadingRequest instance and the progressStatus value of the MLEntityLoadingProcess instance corresponding to all or some models. Correspondingly, the NMS receives the deployment process status of the target model group.
[0291] In another implementation of S805 , the EMS reports the deployment process status of all or part of the models in the target model group to the NMS, including the identifier of the MLEntityLoadingProcess instance and the progressStatus value of the MLEntityLoadingProcess instance corresponding to all or part of the models.
[0292] S806: The NMS sends a control message to the EMS regarding the deployment process of all or some of the models in the target model group. The message includes the identifier of the MLEntityLoadingRequest instance, the identifiers of all or some of the models to be controlled, and the control instructions for the MLEntityLoadingProcess instance corresponding to each model identifier. The EMS executes the control instructions accordingly.
[0293] In another implementation of S806 , the NMS sends a control message of the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance and a control instruction for each MLEntityLoadingProcess instance.
[0294] S807. The EMS sends an indication message to the gNB, indicating activation of the target model group. The indication message includes identification information of the target model group. In response, the gNB receives the indication message.
[0295] S808. The gNB activates the target model group based on the identification information of the target model group.
[0296] Among them, S804, S805, and S806 are all optional steps.
[0297] In the embodiment of the present application, the deployment process of the target model group is managed independently. The NMS can obtain and control the deployment status of each model in the target model group. In the case that some models in the target model group fail to be deployed, the NMS can determine which models failed to be deployed. Based on the identification information of the models that failed to be deployed, these models can be redeployed without redeploying the entire target model group, which is conducive to improving management accuracy and reducing unnecessary resource waste.
[0298] The third case: Scenario 2 + batch management
[0299] FIG9 is a schematic flow chart of a model deployment method 900 provided in an embodiment of the present application. The method 900 includes:
[0300] S901: The NMS sends an MLEntityLoadingPolicy to the EMS, where the MLEntityLoadingPolicy includes identification information of a target model group and a deployment policy of the target model group. Correspondingly, the EMS receives the MLEntityLoadingPolicy.
[0301] Optionally, the MLEntityLoadingPolicy also includes identification information of each model in the target model group.
[0302] In one possible implementation, corresponding to Design 1 above, the MLEntityCoordinationGroup associated with the MLEntityLoadingPolicy can indicate the identification information of the target model group. Alternatively, the value of the memberMLEntityRefList attribute of the MLEntityCoordinationGroup object can be used to indicate the identification information of each model in the target model group. Alternatively, multiple values of the mLEntityId attribute of the MLEntityLoadingPolicy object can be used to indicate the identification information of multiple models in the target model group.
[0303] In another possible implementation, corresponding to the aforementioned second design, the identification information of the target model group may be indicated by the mLEntityGroupInfo property of the MLEntityLoadingPolicy object.
[0304] Optionally, multiple values of the mLEntityId attribute of the MLEntityLoadingPolicy object may be used to indicate identification information of multiple models in the target model group.
[0305] S902: EMS creates an MLEntityLoadingPolicy instance.
[0306] When the EMS determines that the deployment policy of the target model group is satisfied, the EMS executes S903 .
[0307] S903: EMS creates an MLEntityLoadingProcess instance.
[0308] It should be understood that the MLEntityLoadingProcess instance created by the EMS in S903 corresponds to the overall deployment process of the target model group, that is, multiple models in the target model group are deployed and managed as a whole.
[0309] It should also be understood that the deployment process management mode of the target model group can be default or NMS-indicated. In the case of NMS-indicated, the indication information can be indicated by the groupBatchIndicator included in the MLEntityLoadingPolicy. In this embodiment, the groupBatchIndicator indicates that the deployment process management mode of the target model group is batch management.
[0310] S904: The EMS sends the identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance to the NMS. Correspondingly, the NMS receives the identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance.
[0311] Optionally, the EMS may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S902, or may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S903. The identifier of the MLEntityLoadingPolicy instance and the identifier of the MLEntityLoadingProcess instance may be carried in the same message or in different messages, which is not limited in this application.
[0312] S905 : The NMS sends a deployment process status request of the target model group to the EMS, including the identifier of the MLEntityLoadingPolicy instance. Correspondingly, the EMS receives the request.
[0313] In another implementation of S905 , the NMS sends a deployment process status request of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance.
[0314] S906: The EMS reports the deployment process status of the target model group to the NMS, including the identifier of the MLEntityLoadingPolicy instance and the progressStatus value corresponding to the MLEntityLoadingProcess instance. Correspondingly, the NMS receives the deployment process status of the target model group.
[0315] In another implementation of S906 , the EMS reports the deployment process status of the target model group to the NMS, including the identifier of the MLEntityLoadingProcess instance and the progressStatus value corresponding to the MLEntityLoadingProcess instance.
[0316] S907: The NMS sends a control message for the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingPolicy instance and a control instruction. Correspondingly, the EMS executes the control instruction.
[0317] In another implementation of S907 , the NMS sends a control message of the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance and a control instruction.
[0318] S908. The EMS sends an indication message to the gNB, indicating activation of the target model group. The indication message includes identification information of the target model group. In response, the gNB receives the indication message.
[0319] It should be understood that S908 can be understood as the process of the EMS performing model deployment and loading the target model group into the target gNB. Optionally, the number of target gNBs can be one or more, which is not limited in this application.
[0320] Optionally, the indication message may be the second indication message or the third indication message mentioned above, which is not limited in this application.
[0321] S909. The gNB activates the target model group based on the identification information of the target model group.
[0322] In one possible implementation, after S909, method 900 further includes: the gNB performing an inference function based on the target model group.
[0323] It should be understood that the above S905, S906, and S907 are all optional steps.
[0324] The beneficial effects of the embodiment of the present application are similar to those of the aforementioned method 700 and will not be repeated here.
[0325] The fourth case: Scenario 2 + independent management
[0326] FIG10 is a schematic flow chart of a model deployment method 1000 provided in an embodiment of the present application. The method 1000 includes:
[0327] S1001: The NMS sends an MLEntityLoadingPolicy to the EMS, where the MLEntityLoadingPolicy includes identification information of a target model group and a deployment policy of the target model group. Correspondingly, the EMS receives the MLEntityLoadingPolicy.
[0328] Optionally, the MLEntityLoadingPolicy also includes identification information of each model in the target model group.
[0329] It should be understood that the representation of the identification information of the target model group and the identification information of each model in the target model group in method 1000 may be the same as that of the above-mentioned method 900, and will not be repeated here.
[0330] S1002. EMS creates an MLEntityLoadingPolicy instance.
[0331] When the EMS determines that the deployment policy of the target model group is satisfied, the EMS executes S1003 .
[0332] S1003. EMS creates multiple MLEntityLoadingProcess instances.
[0333] It should be understood that the multiple MLEntityLoadingProcess instances correspond one-to-one to the multiple models in the target model group.
[0334] It should also be understood that the deployment process management mode of the target model group can be default or NMS-indicated. In the case of NMS-indicated, the indication information can be indicated by the groupBatchIndicator included in the MLEntityLoadingPolicy. In this embodiment, the groupBatchIndicator indicates that the deployment process management mode of the target model group is independent management.
[0335] S1004: The EMS sends the identifier of the MLEntityLoadingPolicy instance and multiple identifiers corresponding to multiple MLEntityLoadingProcess instances to the NMS. Correspondingly, the NMS receives the message.
[0336] Optionally, the EMS may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S1002, or may send the identifier of the MLEntityLoadingPolicy instance to the NMS after executing S1003. The identifier of the MLEntityLoadingPolicy instance and multiple identifiers corresponding to multiple MLEntityLoadingProcess instances may be carried in the same message or in different messages, which is not limited in this application.
[0337] S1005 . The NMS sends a deployment process status request for all or part of the models in the target model group to the EMS, including the identifier of the MLEntityLoadingPolicy instance and the identification information of all or part of the requested models. Correspondingly, the EMS receives the request.
[0338] In another implementation of S1005 , the NMS sends a deployment process status request of the target model group to the EMS, including all or part of the identification information of the MLEntityLoadingProcess instance.
[0339] S1006: The EMS reports the deployment process status of all or some models in the target model group to the NMS, including the identifier of the MLEntityLoadingPolicy instance and the progressStatus value of the MLEntityLoadingProcess instance corresponding to all or some models. Correspondingly, the NMS receives the deployment process status of the target model group.
[0340] In another implementation of S1006 , the EMS reports the deployment process status of all or some models in the target model group to the NMS, including the identifier of the MLEntityLoadingProcess instance and the progressStatus value of the MLEntityLoadingProcess instance corresponding to all or some models.
[0341] S1007: The NMS sends a control message to the EMS regarding the deployment process of all or some of the models in the target model group. The message includes the identifier of the MLEntityLoadingPolicy instance, the identifiers of all or some of the models to be controlled, and the control instructions for the MLEntityLoadingProcess instance corresponding to each model identifier. The EMS executes the control instructions accordingly.
[0342] In another implementation of S1007 , the NMS sends a control message of the deployment process of the target model group to the EMS, including the identifier of the MLEntityLoadingProcess instance and a control instruction for each MLEntityLoadingProcess instance.
[0343] S1008. The EMS sends an indication message to the gNB, indicating activation of the target model group. The indication message includes identification information of the target model group. In response, the gNB receives the indication message.
[0344] S1009. The gNB activates the target model group based on the identification information of the target model group.
[0345] Among them, S1005, S1006, and S1007 are all optional steps.
[0346] The beneficial effects of the embodiment of the present application are similar to those of the aforementioned method 800 and will not be repeated here.
[0347] It should be understood that the size of the serial numbers of the various steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0348] The model deployment method of an embodiment of the present application is described in detail above in conjunction with Figures 6 to 10. The communication device of an embodiment of the present application will be described in detail below in conjunction with Figures 11 and 12.
[0349] FIG11 shows a communication device 1100 provided in an embodiment of the present application. The communication device 1100 includes a transceiver module 1101 and a processing module 1102 .
[0350] In a possible implementation, the communication device 1100 is used to implement the steps and processes corresponding to the above-mentioned second device.
[0351] Among them, the transceiver module 1101 is used to: receive the identification information of the target model group, the target model group includes multiple models, and the multiple models are jointly trained and / or jointly tested; the processing module 1102 is used to: deploy the target model group based on the identification information of the target model group.
[0352] Optionally, the transceiver module 1101 is also used to: send first identification information and identification information of the first process, the first identification information is associated with the message carrying the identification information of the target model group, there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model group.
[0353] Optionally, the transceiver module 1101 is further used to: receive first indication information, where the first indication information is used to indicate a deployment process management method for the target model group.
[0354] Optionally, the deployment process management mode of the target model group is batch management.
[0355] Optionally, the deployment process management mode of the target model group is batch management by default.
[0356] Optionally, the first process corresponds to the overall deployment process of the target model group.
[0357] Optionally, the transceiver module 1101 is also used to: send first information, the first information includes first identification information and the status of the first process, or the identification information of the first process and the status of the first process, the status including: running, canceling, canceled, paused or ended.
[0358] Optionally, the transceiver module 1101 is also used to: receive second information, the second information is used to control the overall deployment process of the target model group, the second information includes first identification information and control instructions for the first process, or identification information and control instructions of the first process, the control instructions include cancel, pause or start.
[0359] Optionally, the deployment process management mode of the target model group is set to independent management by default.
[0360] Optionally, the deployment process management mode of the target model group is independent management.
[0361] Optionally, the first process includes multiple processes, the multiple processes correspond one-to-one to the identification information of each model in the multiple models, the multiple processes correspond to the deployment processes of the multiple models respectively, and the identification information of the first process includes multiple identification information corresponding to the multiple processes.
[0362] Optionally, the transceiver module 1101 is further configured to: send third information, the third information including the first identification information, identification information of each model corresponding to all or part of the multiple processes, and respective states of all or part of the processes; or,
[0363] Identification information of all or part of the processes, and the status of all or part of the processes, where the status includes: running, canceling, canceled, paused, or ended.
[0364] Optionally, the transceiver module 1101 is also used to: receive fourth information, the fourth information is used to control the deployment process of all or part of the models in the multiple models, the fourth information includes the first identification information, the identification information of each model in all or part of the models, and the control instructions for the deployment process of each model in all or part of the models; or, the identification information of the deployment process of each model in all or part of the models, and the control instructions.
[0365] Optionally, the identification information of the target model group is sent through a request message or a configuration message.
[0366] Optionally, the transceiver module 1101 is further configured to receive identification information of each model in the target model group.
[0367] In another possible implementation, the communication device 1100 is used to implement the steps and processes corresponding to the above-mentioned first device.
[0368] Among them, the processing module 1102 is used to: determine the identification information of the target model group, the target model group includes multiple models, and the multiple models are jointly trained and / or jointly tested; the transceiver module 1101 is used to: send the identification information of the target model group.
[0369] Optionally, the transceiver module 1101 is also used to: receive first identification information and identification information of the first process, the first identification information is associated with the message carrying the identification information of the target model group, there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model group.
[0370] Optionally, the transceiver module 1101 is further used to: send first indication information, where the first indication information is used to indicate a deployment process management method for the target model group.
[0371] Optionally, the deployment process management mode of the target model group is batch management.
[0372] Optionally, the deployment process management mode of the target model group defaults to batch management.
[0373] Optionally, the first process corresponds to the overall deployment process of the target model group.
[0374] Optionally, the transceiver module 1101 is also used to: receive first information, the first information includes first identification information and the status of the first process, or the identification information of the first process and the status of the first process, the status including: running, canceling, canceled, paused or ended.
[0375] Optionally, the transceiver module 1101 is also used to: send second information, the second information is used to control the overall deployment process of the target model group, the second information includes first identification information and control instructions for the first process, or identification information and control instructions of the first process, the control instructions include cancel, pause or start.
[0376] Optionally, the deployment process management mode of the target model group is set to independent management by default.
[0377] Optionally, the deployment process management mode of the target model group is independent management.
[0378] Optionally, the first process includes multiple processes, the multiple processes correspond one-to-one to the identification information of each model in the multiple models, the multiple processes correspond to the deployment processes of the multiple models respectively, and the identification information of the first process includes multiple identification information corresponding to the multiple processes.
[0379] Optionally, the transceiver module 1101 is also used to: receive third information, the third information including the first identification information, the identification information of each model corresponding to all or part of the multiple processes, and the respective status of all or part of the processes; or, the identification information of all or part of the processes, and the respective status of all or part of the processes, the status including: running, canceling, canceled, paused or ended.
[0380] Optionally, the transceiver module 1101 is also used to: send fourth information, the fourth information is used to control the deployment process of all or part of the models in the multiple models, the fourth information includes the first identification information, the identification information of each model in all or part of the models, and the control instructions for the deployment process of each model in all or part of the models; or, the identification information of the deployment process of each model in all or part of the models, and the control instructions.
[0381] Optionally, the identification information of the target model group is sent through a request message or a configuration message.
[0382] Optionally, the transceiver module 1101 is further configured to send identification information of each model in the target model group.
[0383] It should be understood that the device 1100 here is embodied in the form of a functional module. The term "module" here can refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combined logic circuit and / or other suitable components that support the described functions. In an optional example, those skilled in the art will understand that the device 1100 can be specifically the first device or the second device in the above-mentioned embodiment, or the functions described in the above-mentioned embodiment can be integrated in the device 1100, and the device 1100 can be used to execute the various processes and / or steps corresponding to the first device or the second device in the above-mentioned method embodiment. To avoid repetition, they will not be described here.
[0384] The apparatus 1100 has the function of implementing the corresponding steps performed by the first apparatus or the second apparatus in the above method. The above functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0385] In an embodiment of the present application, the device 1100 in FIG. 11 may also be a chip or a chip system, such as a system on chip (SoC).
[0386] Figure 12 shows a schematic block diagram of a communication device 1200 provided in an embodiment of the present application. The device 1200 includes a processor 1201, a transceiver 1202, and a memory 1203. The processor 1201, the transceiver 1202, and the memory 1203 communicate with each other via an internal connection path. The memory 1203 is used to store instructions, and the processor 1201 is used to execute the instructions stored in the memory 1203 to control the transceiver 1202 to send and / or receive signals.
[0387] It should be understood that apparatus 1200 may be specifically the first apparatus or the second apparatus in the aforementioned embodiment, and may be used to execute the various steps and / or processes corresponding to the first apparatus or the second apparatus in the aforementioned method embodiment. Optionally, the memory 1203 may include read-only memory and random access memory, and provide instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store device type information. The processor 1201 may be used to execute instructions stored in the memory, and when the processor 1201 executes the instructions stored in the memory, the processor 1201 is used to execute the various steps and / or processes of the aforementioned method embodiment. The transceiver 1202 may include a transmitter and a receiver. The transmitter may be used to implement the various steps and / or processes corresponding to the aforementioned transceiver for performing a sending action, and the receiver may be used to implement the various steps and / or processes corresponding to the aforementioned transceiver for performing a receiving action.
[0388] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0389] During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor executes the instructions in the memory, and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.
[0390] The present application also provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to implement the method shown in the above method embodiment.
[0391] The present application also provides a computer program product, which includes computer program code (also referred to as a computer program or instruction). When the computer program runs on a computer, the computer can execute the method shown in the above method embodiment.
[0392] The present application also provides a communication method, which is applied to a communication system including a first device and a second device, and the method shown in the above method embodiment is executed by the second device and / or the second device.
[0393] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0394] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0395] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0396] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0397] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0398] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program codes.< / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc> < / ioc>
Claims
1. A model deployment method, characterized in that: include: Receiving identification information of a target model group, the target model group including a plurality of models, the plurality of models being jointly trained and / or jointly tested; The target model group is deployed based on the identification information of the target model group.
2. The method according to claim 1, characterized in that The method further comprises: Send first identification information and identification information of a first process, wherein the first identification information is associated with a message carrying the identification information of the target model group, and there is a mapping relationship between the first identification information and the first process, and the first process is used to manage the deployment process of the target model group.
3. The method according to claim 2, characterized in that The deployment process management method of the target model group is batch management.
4. The method according to claim 3, characterized in that The first process corresponds to the overall deployment process of the target model group.
5. The method according to claim 3 or 4, characterized in that: The method further comprises: Sending first information, wherein the first information includes the first identification information and a state of the first process, or the identification information of the first process and a state of the first process, wherein the state includes: running, canceling, canceled, paused, or ended.
6. The method according to any one of claims 3 to 5, characterized in that The method further comprises: Receive second information, where the second information is used to control the overall deployment process of the target model group, and the second information includes the first identification information and control instructions for the first process, or the identification information of the first process and the control instructions, and the control instructions include canceling, pausing or starting.
7. The method according to claim 2, characterized in that The deployment process management mode of the target model group is independent management.
8. The method according to claim 7, characterized in that The first process includes multiple processes, and the multiple processes correspond one-to-one to identification information of each model in the multiple models. The multiple processes correspond to deployment processes of the multiple models respectively, and the identification information of the first process includes multiple identification information corresponding to the multiple processes.
9. The method according to claim 7 or 8, characterized in that: The method further comprises: sending third information, wherein the third information includes the first identification information, identification information of each model corresponding to all or part of the multiple processes, and states of all or part of the processes; or, The identification information of all or part of the processes, and the status of all or part of the processes, wherein the status includes: running, canceling, canceled, paused or ended.
10. The method according to any one of claims 7 to 9, characterized in that The method further comprises: receiving fourth information, the fourth information being used to control the deployment process of all or part of the multiple models, the fourth information including the first identification information, identification information of each model in the all or part of the models, and a control instruction for the deployment process of each model in the all or part of the models; or The identification information of the deployment process of each model in the whole or part of the models, and the control instructions.
11. The method according to any one of claims 1 to 10, characterized in that The method further comprises: First indication information is received, where the first indication information is used to indicate a deployment process management method for the target model group.
12. The method according to any one of claims 3 to 6, characterized in that The deployment process management mode of the target model group is batch management by default.
13. The method according to any one of claims 7 to 10, characterized in that The deployment process management mode of the target model group is independent management by default.
14. The method according to any one of claims 1 to 13, characterized in that The identification information of the target model group is sent via a request message or a configuration message.
15. The method according to any one of claims 1 to 14, characterized in that The method further comprises: The identification information of each model in the target model group is received.
16. A model deployment method, characterized in that: include: Determine identification information of a target model group, where the target model group includes a plurality of models, and the plurality of models are jointly trained and / or jointly tested; The identification information of the target model group is sent.
17. The method according to claim 16, characterized in that The method further comprises: Receive first identification information and identification information of a first process, wherein the first identification information is associated with a message carrying the identification information of the target model group, and a mapping relationship exists between the first identification information and the first process, and the first process is used to manage the deployment process of the target model group.
18. The method according to claim 17, characterized in that The deployment process management method of the target model group is batch management.
19. The method according to claim 18, characterized in that The first process corresponds to the overall deployment process of the target model group.
20. The method according to claim 18 or 19, characterized in that The method further comprises: Receive first information, where the first information includes the first identification information and a state of the first process, or the identification information of the first process and a state of the first process, where the state includes: running, canceling, canceled, paused, or ended.
21. The method according to any one of claims 17 to 19, characterized in that The method further comprises: Sending second information, where the second information is used to control the overall deployment process of the target model group, the second information includes the first identification information and control instructions for the first process, or the identification information of the first process and the control instructions, and the control instructions include canceling, pausing or starting.
22. The method according to claim 17, characterized in that The deployment process management mode of the target model group is independent management.
23. The method according to claim 22, characterized in that The first process includes multiple processes, and the multiple processes correspond one-to-one to identification information of each model in the multiple models. The multiple processes correspond to deployment processes of the multiple models respectively, and the identification information of the first process includes multiple identification information corresponding to the multiple processes.
24. The method according to claim 22 or 23, characterized in that The method further comprises: receiving third information, the third information including the first identification information, identification information of each model corresponding to all or part of the multiple processes, and states of all or part of the processes; or, The identification information of all or part of the processes, and the status of all or part of the processes, wherein the status includes: running, canceling, canceled, paused or ended.
25. The method according to any one of claims 22 to 24, characterized in that The method further comprises: sending fourth information, where the fourth information is used to control the deployment process of all or part of the multiple models, and the fourth information includes the first identification information, identification information of each model in all or part of the models, and a control instruction for the deployment process of each model in all or part of the models; or The identification information of the deployment process of each model in the whole or part of the models, and the control instructions.
26. The method according to any one of claims 16 to 25, characterized in that The method further comprises: Sending first indication information, where the first indication information is used to indicate a deployment process management method for the target model group.
27. The method according to any one of claims 18 to 21, characterized in that The deployment process management mode of the target model group is batch management by default.
28. The method according to any one of claims 22 to 25, characterized in that The deployment process management mode of the target model group is independent management by default.
29. The method according to any one of claims 16 to 28, characterized in that The identification information of the target model group is sent via a request message or a configuration message.
30. The method according to any one of claims 16 to 29, characterized in that The method further comprises: The identification information of each model in the target model group is sent.
31. A communication device, characterized in that: include: A module for executing the method of any one of claims 1 to 15, or a module for executing the method of any one of claims 16 to 30.
32. A communication device, characterized in that: include: A processor, wherein the processor is coupled to a memory, the memory stores computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory, so that the processor executes the method according to any one of claims 1 to 15, or executes the method according to any one of claims 16 to 30.
33. A computer-readable storage medium, characterized in that: Used to store a computer program, the computer program comprising instructions for implementing the method according to any one of claims 1 to 15, or instructions for executing the method according to any one of claims 16 to 30.
34. A computer program product, characterized in that The computer program product includes computer program codes, and when the computer program codes are run on a computer, the computer is enabled to implement the method according to any one of claims 1 to 15, or to execute the method according to any one of claims 16 to 30.
35. A communication system, characterized in that: include: An apparatus for performing the method according to any one of claims 1 to 15, and / or an apparatus for performing the method according to any one of claims 16 to 30.
36. A communication method, characterized in that: Applied to a communication system including a first device and a second device, the method comprises: executing the method according to any one of claims 1 to 15 by the second device, and executing the method according to any one of claims 16 to 30 by the first device.
Citation Information
Patent Citations
Model deployment method and communication device
CN120200902A
AI model deployment method and system and storage medium
CN115392332A
Multi-model deployment reasoning method and system, network equipment and readable storage medium
CN116977831A
Model group training method and associated model group deployment method and device
CN117217298A
Model deployment method, model deployment device and terminal equipment
US20220100486A1